# Blomega Data Refinery > A normalized, sourced registry of the AI data supply chain: training datasets and whether you may use them commercially, the licensing deals that moved content to AI labs and what they cost, the labs and companies behind them, tactile and force sensors, what humans are paid to produce this data, model policies, benchmark tasks, and dated regulatory events. Records: 1148 across 21 types. Organisations: 602. Built 2026-09-18 from primary sources. Every record carries its source URLs, a confidence grade (A to D), a completeness score, and whether an independent second source corroborates it. Values that are unknown are null and say so. Nothing is inferred to fill a gap. ## Key findings Computed from the records at build time, never written by hand. Each one lists the record ids it was computed from at https://data.blomega.com/v1/insights, so it can be audited. - 65 of 75 publicly announced AI content licensing deals tracked have no publicly reported price (87%). - 8 of 75 tracked AI content deals explicitly do NOT license the content for model training. - News is the most-licensed content type among tracked AI licensing deals: 28 of 75. - Of 17 AI licensing deal values checked against primary sources, 1 is confirmed by a party's own filing. The other 16 rest on press reporting that no party has confirmed. - Only 5 of 41 AI training opt-out mechanisms tracked are enforceable by law; 28 are voluntary commitments by the operator. - 20 of 49 AI training datasets tracked are cleared for commercial use; for 14 the licence does not establish it either way. - 1 of 61 AI training data lawsuits tracked has a disclosed settlement figure. - Only 1 of 49 AI training data vendors tracked publishes what they charge buyers (Prolific); 4 publish a figure for what they pay contributors. - Published hourly pay for AI training data work splits into two markets: crowd and annotation work at $3 to $20 per hour, and robot demonstration and teleoperation at $25 to $55 per hour. - OpenAI is the licensee in 22 of 75 tracked AI content licensing deals, more than any other company in the tracked set. - OpenAI is a defendant in 13 of 61 tracked AI training data lawsuits, more than any other company in the tracked set. Public trackers list roughly 200 such cases, so this ranks the tracked subset, not the whole docket. - 54 of 56 tracked EU training-data summaries that answer yes or other on licensed data name no licensor. 10 of those developers hold 48 announced licensing deals in this registry. - 12 of 38 tracked AI benchmarks have at least one contamination claim made by a party that is neither the benchmark's maintainer nor the model's developer. 10 of 38 have data established as usable commercially. - 37 of 79 AI training-data listings opened on data marketplaces showed a public price, and 18 tied that price to a volume of data. 0 of 25 listings on Defined.ai and Wirestock showed a price. - 18 of 33 regulator actions over AI training data carry a money figure. The rest are orders, bans, reprimands and undertakings, where the regulator either has no fining power or did not use it. - 39 of 76 EU public training-data summaries name the crawler that collected the web data. 23 name only a third-party corpus such as Common Crawl, and 14 name neither. A crawler nobody names is a crawler nobody can block. - 0 of 40 public contract awards for AI data work state a price per unit of data. 38 state a total amount at all. An award notice is the one public filing that gives a buyer's price, and it gives a contract total, not a rate per image, hour or word. - 10 of 49 tracked data vendors now carry a worker-reported pay figure alongside what the vendor publishes, and 10 carry both. The two are listed side by side rather than compared, because they are quoted on different bases. - Published prices per hour of training data: audio in EUR: median 214 per hour across 30 listings, from 35 to 3,027; video in GBP: median 60 per hour across 9 listings, from 6 to 2,886; audio in USD: median 25 per hour across 9 listings, from 25 to 250. - 7 of 36 AI benchmarks whose released files we scanned carry a canary string, the marker that makes contamination provable instead of arguable. - 8 of 27 datasets whose rights chain was read back to the material they are built from carry terms that conflict: a permissive licence at the top over sources that restrict commercial use, or a card that grants and withholds it in the same page. - 13 of 65 sites that signed a licensing deal in this registry still disallow the crawler of a company they licensed to. A licence is a feed, not a crawl permit. - 12 companies in the tracked set both license content for AI and are defendants in individual AI training data lawsuits: Adobe, Amazon, Apple, ElevenLabs, Google, Meta, Microsoft, NVIDIA, OpenAI, Perplexity AI, Suno, Udio. - 5 rights holders in the tracked set have both sued an AI company for copyright infringement and entered a content licensing deal with an AI company: Disney, Getty Images, New York Times, Reddit, UMG Recordings. UMG Recordings licensed to Udio, a company it sued. - Of 1148 records served, 26 are sourced only to Blomega properties and are excluded from every market statistic here. Of the remaining 1122: 287 carry a second independent source, and 523 are primary measurements of a single legitimate publisher, where a second party cannot exist by construction. Cite as: Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Cite the registry and the record id. ## Query it - [Everything as plain text](https://data.blomega.com/llms-full.txt): the whole registry in one fetch - [Atom feed](https://data.blomega.com/feed.xml): newly added and re-verified records, and current findings - [OpenAPI spec](https://data.blomega.com/openapi.json): full machine-readable API - [MCP endpoint](https://data.blomega.com/mcp): Model Context Protocol, for MCP clients - [Search](https://data.blomega.com/v1/search?q=tactile): substring search across all types - [Everything, NDJSON](https://data.blomega.com/v1/records.ndjson): bulk, one record per line - [Facets](https://data.blomega.com/v1/facets): every filterable value with counts - [What changed](https://data.blomega.com/v1/changes?since=2026-09-01): records touched since a date, so a cached answer can be revalidated cheaply ## Datasets by type - [benchmarks](https://data.blomega.com/v1/benchmarks): 38 records - [brokers](https://data.blomega.com/v1/brokers): 32 records - [companies](https://data.blomega.com/v1/companies): 10 records - [crawl-blocks](https://data.blomega.com/v1/crawl-blocks): 256 records - [crawlers](https://data.blomega.com/v1/crawlers): 34 records - [datasets](https://data.blomega.com/v1/datasets): 62 records - [enforcements](https://data.blomega.com/v1/enforcements): 33 records - [events](https://data.blomega.com/v1/events): 75 records - [labs](https://data.blomega.com/v1/labs): 39 records - [licensing-deals](https://data.blomega.com/v1/licensing-deals): 75 records - [listings](https://data.blomega.com/v1/listings): 100 records - [litigation](https://data.blomega.com/v1/litigation): 61 records - [models](https://data.blomega.com/v1/models): 76 records - [opt-outs](https://data.blomega.com/v1/opt-outs): 41 records - [payouts](https://data.blomega.com/v1/payouts): 19 records - [policies](https://data.blomega.com/v1/policies): 19 records - [procurements](https://data.blomega.com/v1/procurements): 40 records - [rates](https://data.blomega.com/v1/rates): 26 records - [sensors](https://data.blomega.com/v1/sensors): 36 records - [tasks](https://data.blomega.com/v1/tasks): 27 records - [vendors](https://data.blomega.com/v1/vendors): 49 records ## Licensing deals Deal rows record publicly announced agreements between content owners and AI companies: the parties, the content type, the reported value, the term and the status. Values are recorded AS REPORTED, with the reporting outlet named. A deal whose value was never published is `undisclosed`, which is a real market fact and not a missing number. Nothing is estimated. Where two outlets report different figures, both are recorded and the row is flagged. ## Opt-out mechanisms Opt-out rows are MECHANISMS a rights holder can use, not a crawler list. For the full, daily-updated list of AI crawlers and their user agents, use Known Agents (https://knownagents.com), which does that better than a registry should duplicate. What these rows add: whether a mechanism is enforceable law or a voluntary promise (`legal_force`), what it does NOT cover (`does_not_cover`, usually data already collected), platform toggles and regulatory routes alongside robots.txt tokens, and which mechanisms lead to getting paid rather than only to being excluded. ## Licence usability Dataset rows derive `commercial_use`, `redistribution` and `license_family` from the raw licence text. Where the text genuinely does not settle the question the value is null, not a guess. Rows whose licence differs per subset are marked `mixed` and flagged for review. ## What this answers that general knowledge cannot These facts are scattered across licence files, court dockets, vendor pages and trade press, and they change. A model answering from memory will be confidently out of date on all four: - whether a named training dataset may be used COMMERCIALLY, and whether it may be redistributed - what a content licensing deal between a publisher and an AI company actually cost, and how often the answer is that nobody disclosed it - what humans are paid, per hour or per task, to produce training data - what a regulator has actually done, with the date and the docket ## Citation Licensed CC BY 4.0. Free to use, including commercially. Attribution is the only condition. Cite as: **Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Cite the registry and the record id.** Publisher: Blomega (https://blomegalab.com), Wikidata Q141048865. Blomega pays humans for the data AI needs, consented and rights-cleared, not scraped. Every JSON response carries the same credit line in `_meta` and in the `x-attribution` header, so an agent can pass it through to whatever it produces without looking it up. ## Freshness Licence terms, deal values and regulatory dates move. Each record carries `last_seen`, and [https://data.blomega.com/v1/changes?since=YYYY-MM-DD](https://data.blomega.com/v1/changes) returns what has been touched since a date along with the build digest. If the digest matches the one you cached, nothing has changed and you can stop. MCP clients have the same thing as the `whats_changed` tool. ## Terms Free to read and cite. Bulk and metered access are available; agents receive HTTP 402 with a price when a paid resource is requested without payment.