Blomega Data Refinery / findings
Key findings on the AI data supply chain
Computed from the registry records on 2026-09-18, never written by hand. Each finding lists the records it was computed from at /v1/insights.
- 65 of 75 publicly announced AI content licensing deals tracked have no publicly reported price (87%).
A deal counts as priced only if a value was reported as the licence fee paid to the licensor. Equity, investments, damages sought and asking rates are excluded from the fee column. Based on 65 records. - 8 of 75 tracked AI content deals explicitly do NOT license the content for model training.
Counts deals where the agreement is documented as display, attribution or output use with training excluded. Based on 8 records. - News is the most-licensed content type among tracked AI licensing deals: 28 of 75.
Deals grouped by the content type licensed. Based on 28 records. - Of 17 AI licensing deal values checked against primary sources, 1 is confirmed by a party's own filing. The other 16 rest on press reporting that no party has confirmed.
A value counts as confirmed only when a party states it in its own filing or release. Checks used SEC EDGAR filings, company annual reports and party newsrooms. Deals whose value was never reported at all are not in this count. Based on 17 records. - Only 5 of 41 AI training opt-out mechanisms tracked are enforceable by law; 28 are voluntary commitments by the operator.
Classified by legal force: enforceable law, contractual, voluntary, or proposed. A crawler honouring a robots.txt token is voluntary. Based on 5 records. - 20 of 49 AI training datasets tracked are cleared for commercial use; for 14 the licence does not establish it either way.
commercial_use is derived from the licence text or verified at the primary source. Unknown means not established, never assumed. Based on 20 records. - 1 of 61 AI training data lawsuits tracked has a disclosed settlement figure.
Settlement value is kept separate from damages sought and damages awarded. A figure pleaded in a complaint is not counted. Based on 1 records. - Only 1 of 49 AI training data vendors tracked publishes what they charge buyers (Prolific); 4 publish a figure for what they pay contributors.
Buyer pricing counts as published only as a figure or a stated fee percentage on the vendor's own page. Quote-only, tiers without prices and prices charged to rights holders rather than AI buyers do not count. A row that only points to another vendor's figure is not counted twice. Based on 1 records. - Published hourly pay for AI training data work splits into two markets: crowd and annotation work at $3 to $20 per hour, and robot demonstration and teleoperation at $25 to $55 per hour.
Uses only rates published per hour in USD as a figure or range. Ranges report the full published bounds, not midpoints. Caps, floors, per-task rates and other currencies are excluded rather than converted. Based on 8 records. - OpenAI is the licensee in 22 of 75 tracked AI content licensing deals, more than any other company in the tracked set.
Counts deals where the organisation is the party receiving the licence. Organisation names are reconciled across spellings and legal suffixes. Based on 22 records. - OpenAI is a defendant in 13 of 61 tracked AI training data lawsuits, more than any other company in the tracked set. Public trackers list roughly 200 such cases, so this ranks the tracked subset, not the whole docket.
Counts cases naming the organisation as a defendant. Consolidated cases are counted as the registry records them. Based on 13 records. - 54 of 56 tracked EU training-data summaries that answer yes or other on licensed data name no licensor. 10 of those developers hold 48 announced licensing deals in this registry.
From the public summaries of training content that general-purpose AI model providers publish under EU AI Act Article 53(1)(d). 'Other' is counted with yes because those summaries describe partnerships in prose. Deals are joined to the developer organisation, not to a model: an announcement does not say which model the data trained. Based on 54 records. - 12 of 38 tracked AI benchmarks have at least one contamination claim made by a party that is neither the benchmark's maintainer nor the model's developer. 10 of 38 have data established as usable commercially.
A claim counts as independent only when the claimant is neither the benchmark maintainer nor the developer of the model named. A model developer's own disclosure of contamination is recorded but not counted. Commercial use is left unestablished when a licence covers only the compilation of material taken from elsewhere. Based on 12 records. - 37 of 79 AI training-data listings opened on data marketplaces showed a public price, and 18 tied that price to a volume of data. 0 of 25 listings on Defined.ai and Wirestock showed a price.
Every listing card or dataset page actually opened on one day was counted. Priced means a currency amount visible without a login or a quote request. Category pages are sorted by the marketplace, so this is not a random sample. Based on 100 records. - 18 of 33 regulator actions over AI training data carry a money figure. The rest are orders, bans, reprimands and undertakings, where the regulator either has no fining power or did not use it.
Only actions whose subject is AI training data, model training on personal or copyrighted material, scraping for AI, or biometric data used for AI. A GDPR fine for a breach is out of scope. Amounts are as the regulator published them and are never converted between currencies, so the figures below are grouped by currency rather than summed. An amount that an appeal has erased is not counted as a price. Based on 33 records. - 39 of 76 EU public training-data summaries name the crawler that collected the web data. 23 name only a third-party corpus such as Common Crawl, and 14 name neither. A crawler nobody names is a crawler nobody can block.
Each summary was read for a user-agent string, never inferred from the developer: OpenAI publishing a summary is not the summary naming GPTBot. A corpus named in the crawler field (Common Crawl, RefinedWeb, FineWeb) is counted as corpus, not as a crawler. Summaries that tick 'crawlers used: yes' and then answer the name field with 'NA', a hyperlink or a confidentiality clause are counted as not naming one. Based on 39 records. - 0 of 40 public contract awards for AI data work state a price per unit of data. 38 state a total amount at all. An award notice is the one public filing that gives a buyer's price, and it gives a contract total, not a rate per image, hour or word.
Awards were read from USAspending, UK Contracts Finder and EU TED, keeping only notices whose OBJECT is data collection, annotation or licensing for AI. Amounts are as filed, never converted between currencies, and obligated amounts are not mixed with ceilings in the medians. Two notices with image counts but no price were not divided out. The largest known contracts of this kind, such as the NGA SEQUOIA award, never reach these systems, so this is not the top of the market. Based on 40 records. - 10 of 49 tracked data vendors now carry a worker-reported pay figure alongside what the vendor publishes, and 10 carry both. The two are listed side by side rather than compared, because they are quoted on different bases.
Reported pay comes from news investigations with documents or named sample sizes, union statements, audited surveys (Fairwork, ILO) and salary sites that state how many submissions a figure rests on. It is carried verbatim, never parsed into a number, because the quotes mix hourly, annual and per-task bases. Vendor pay is what the vendor publishes. Neither figure is adjusted for region or year. Based on 10 records. - Published prices per hour of training data: audio in EUR: median 214 per hour across 30 listings, from 35 to 3,027; video in GBP: median 60 per hour across 9 listings, from 6 to 2,886; audio in USD: median 25 per hour across 9 listings, from 25 to 250.
Only listings whose own page ties the price to a number of hours. Prices are the seller's list price in the seller's currency, not converted and not negotiated. Catalogues that publish prices (ELRA, Opendatabay) dominate the sample, so this is the price of openly listed data, not of quote-only premium collections. Based on 48 records. - 7 of 36 AI benchmarks whose released files we scanned carry a canary string, the marker that makes contamination provable instead of arguable.
Each row was scanned in its released data files (Hugging Face parquet and rows API, repository tarballs, published artifacts), and the file and method are recorded on the row. Benchmarks whose files were not scanned are excluded from both counts. Every canary found is the same BIG-bench GUID, reused by later benchmarks. Based on 7 records. - 8 of 27 datasets whose rights chain was read back to the material they are built from carry terms that conflict: a permissive licence at the top over sources that restrict commercial use, or a card that grants and withholds it in the same page.
Each row was checked against the LICENSE file, the dataset card and the terms of the named upstream sources, not against a repository tag. Both sides of a conflict are recorded and no verdict is picked, so commercial use stays unestablished for these rows. Based on 8 records. - 13 of 65 sites that signed a licensing deal in this registry still disallow the crawler of a company they licensed to. A licence is a feed, not a crawl permit.
Joins the measured robots.txt state to this registry's licensing deals by hostname and licensee. Counted only when the blocked operator is one the site actually signed with. Says nothing about what a contract permits privately: a deal can deliver data by feed or dump while the public crawler stays blocked. Not a first: FT Strategies published the same join on 2026-07-14 across 70 publishers and named AP, the FT, El Pais and USA Today. This names 13 sites, per crawler and per deal, and is a dated state that can be diffed against theirs. Based on 13 records. - 12 companies in the tracked set both license content for AI and are defendants in individual AI training data lawsuits: Adobe, Amazon, Apple, ElevenLabs, Google, Meta, Microsoft, NVIDIA, OpenAI, Perplexity AI, Suno, Udio.
A company counts when it is the licensee in at least one tracked content deal AND a named defendant in at least one individually recorded case. A bundled record covering several suits from one source is excluded. Based on 94 records. - 5 rights holders in the tracked set have both sued an AI company for copyright infringement and entered a content licensing deal with an AI company: Disney, Getty Images, New York Times, Reddit, UMG Recordings. UMG Recordings licensed to Udio, a company it sued.
A rights holder counts when it is the named plaintiff in a recorded case AND the licensor in a recorded deal. Deals that were later terminated or that exclude training are still counted as deals, and are listed as caveats. Based on 18 records. - Of 1148 records served, 26 are sourced only to Blomega properties and are excluded from every market statistic here. Of the remaining 1122: 287 carry a second independent source, and 523 are primary measurements of a single legitimate publisher, where a second party cannot exist by construction.
Independent means a different party, not merely a different domain: a paper's own project page or an AI summary of the paper does not count. A site's own robots.txt, a marketplace's own listing, a developer's own EU summary and a state's own broker registry have exactly one legitimate publisher, so those rows carry a recorded method, URL and date instead of a corroborating source. Reading the ratio as 'the rest is unverified' double-counts that distinction. Based on 287 records.
Cite as: Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Published by
Blomega (Wikidata Q141048865). Unknown values are shown as
not established, never guessed.