Blomega Data Refinery
The AI data supply chain, sourced. 1148 records across 21 types, built 2026-09-18. Every record carries its source and a confidence grade; unknown values say so instead of guessing.
Key findings
- 65 of 75 publicly announced AI content licensing deals tracked have no publicly reported price (87%).
- 8 of 75 tracked AI content deals explicitly do NOT license the content for model training.
- News is the most-licensed content type among tracked AI licensing deals: 28 of 75.
- Of 17 AI licensing deal values checked against primary sources, 1 is confirmed by a party's own filing. The other 16 rest on press reporting that no party has confirmed.
- Only 5 of 41 AI training opt-out mechanisms tracked are enforceable by law; 28 are voluntary commitments by the operator.
- 20 of 49 AI training datasets tracked are cleared for commercial use; for 14 the licence does not establish it either way.
- 1 of 61 AI training data lawsuits tracked has a disclosed settlement figure.
- Only 1 of 49 AI training data vendors tracked publishes what they charge buyers (Prolific); 4 publish a figure for what they pay contributors.
- Published hourly pay for AI training data work splits into two markets: crowd and annotation work at $3 to $20 per hour, and robot demonstration and teleoperation at $25 to $55 per hour.
- OpenAI is the licensee in 22 of 75 tracked AI content licensing deals, more than any other company in the tracked set.
- OpenAI is a defendant in 13 of 61 tracked AI training data lawsuits, more than any other company in the tracked set. Public trackers list roughly 200 such cases, so this ranks the tracked subset, not the whole docket.
- 54 of 56 tracked EU training-data summaries that answer yes or other on licensed data name no licensor. 10 of those developers hold 48 announced licensing deals in this registry.
- 12 of 38 tracked AI benchmarks have at least one contamination claim made by a party that is neither the benchmark's maintainer nor the model's developer. 10 of 38 have data established as usable commercially.
- 37 of 79 AI training-data listings opened on data marketplaces showed a public price, and 18 tied that price to a volume of data. 0 of 25 listings on Defined.ai and Wirestock showed a price.
- 18 of 33 regulator actions over AI training data carry a money figure. The rest are orders, bans, reprimands and undertakings, where the regulator either has no fining power or did not use it.
- 39 of 76 EU public training-data summaries name the crawler that collected the web data. 23 name only a third-party corpus such as Common Crawl, and 14 name neither. A crawler nobody names is a crawler nobody can block.
- 0 of 40 public contract awards for AI data work state a price per unit of data. 38 state a total amount at all. An award notice is the one public filing that gives a buyer's price, and it gives a contract total, not a rate per image, hour or word.
- 10 of 49 tracked data vendors now carry a worker-reported pay figure alongside what the vendor publishes, and 10 carry both. The two are listed side by side rather than compared, because they are quoted on different bases.
- Published prices per hour of training data: audio in EUR: median 214 per hour across 30 listings, from 35 to 3,027; video in GBP: median 60 per hour across 9 listings, from 6 to 2,886; audio in USD: median 25 per hour across 9 listings, from 25 to 250.
- 7 of 36 AI benchmarks whose released files we scanned carry a canary string, the marker that makes contamination provable instead of arguable.
- 8 of 27 datasets whose rights chain was read back to the material they are built from carry terms that conflict: a permissive licence at the top over sources that restrict commercial use, or a card that grants and withholds it in the same page.
- 13 of 65 sites that signed a licensing deal in this registry still disallow the crawler of a company they licensed to. A licence is a feed, not a crawl permit.
- 12 companies in the tracked set both license content for AI and are defendants in individual AI training data lawsuits: Adobe, Amazon, Apple, ElevenLabs, Google, Meta, Microsoft, NVIDIA, OpenAI, Perplexity AI, Suno, Udio.
- 5 rights holders in the tracked set have both sued an AI company for copyright infringement and entered a content licensing deal with an AI company: Disney, Getty Images, New York Times, Reddit, UMG Recordings. UMG Recordings licensed to Udio, a company it sued.
- Of 1148 records served, 26 are sourced only to Blomega properties and are excluded from every market statistic here. Of the remaining 1122: 287 carry a second independent source, and 523 are primary measurements of a single legitimate publisher, where a second party cannot exist by construction.
All findings, with how each was computed
Registry
For agents
POST /mcp: Model Context Protocol, start withkey_findingsororg_profile- llms.txt and llms-full.txt, the whole registry in one fetch
- OpenAPI, /v1/insights, /v1/changes, /v1/orgs/{name}
Free to use under CC BY 4.0. Cite as Blomega Data Refinery, https://data.blomega.com.
Published by Blomega, which pays humans for the data AI needs, consented and rights-cleared, not scraped.