<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://data.blomega.com/feed.xml</id>
  <title>Blomega Data Refinery: updates</title>
  <subtitle>Newly added and re-verified records in the AI data supply chain registry, and current findings. Blomega Data Refinery, https://data.blomega.com, CC BY 4.0.</subtitle>
  <link href="https://data.blomega.com/feed.xml" rel="self"/>
  <link href="https://data.blomega.com/"/>
  <updated>2026-09-18T00:00:00Z</updated>
  <author><name>Blomega</name><uri>https://blomegalab.com</uri></author>
  <rights>CC BY 4.0. Cite as Blomega Data Refinery, https://data.blomega.com, CC BY 4.0.</rights>
  <entry>
    <id>https://data.blomega.com/findings#deals-undisclosed</id>
    <title>Finding: 65 of 75 publicly announced AI content licensing deals tracked have no publicly reported price (87%).</title>
    <link href="https://data.blomega.com/findings#deals-undisclosed"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>65 of 75 publicly announced AI content licensing deals tracked have no publicly reported price (87%). Method: A deal counts as priced only if a value was reported as the licence fee paid to the licensor. Equity, investments, damages sought and asking rates are excluded from the fee column.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#deals-exclude-training</id>
    <title>Finding: 8 of 75 tracked AI content deals explicitly do NOT license the content for model training.</title>
    <link href="https://data.blomega.com/findings#deals-exclude-training"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>8 of 75 tracked AI content deals explicitly do NOT license the content for model training. Method: Counts deals where the agreement is documented as display, attribution or output use with training excluded.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#deals-top-content-type</id>
    <title>Finding: News is the most-licensed content type among tracked AI licensing deals: 28 of 75.</title>
    <link href="https://data.blomega.com/findings#deals-top-content-type"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>News is the most-licensed content type among tracked AI licensing deals: 28 of 75. Method: Deals grouped by the content type licensed.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#deals-value-provenance</id>
    <title>Finding: Of 17 AI licensing deal values checked against primary sources, 1 is confirmed by a party's own filing. The other 16 rest on press reporting that no party has confirmed.</title>
    <link href="https://data.blomega.com/findings#deals-value-provenance"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>Of 17 AI licensing deal values checked against primary sources, 1 is confirmed by a party's own filing. The other 16 rest on press reporting that no party has confirmed. Method: A value counts as confirmed only when a party states it in its own filing or release. Checks used SEC EDGAR filings, company annual reports and party newsrooms. Deals whose value was never reported at all are not in this count.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#optouts-enforceable</id>
    <title>Finding: Only 5 of 41 AI training opt-out mechanisms tracked are enforceable by law; 28 are voluntary commitments by the operator.</title>
    <link href="https://data.blomega.com/findings#optouts-enforceable"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>Only 5 of 41 AI training opt-out mechanisms tracked are enforceable by law; 28 are voluntary commitments by the operator. Method: Classified by legal force: enforceable law, contractual, voluntary, or proposed. A crawler honouring a robots.txt token is voluntary.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#datasets-commercial</id>
    <title>Finding: 20 of 49 AI training datasets tracked are cleared for commercial use; for 14 the licence does not establish it either way.</title>
    <link href="https://data.blomega.com/findings#datasets-commercial"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>20 of 49 AI training datasets tracked are cleared for commercial use; for 14 the licence does not establish it either way. Method: commercial_use is derived from the licence text or verified at the primary source. Unknown means not established, never assumed.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#litigation-settled-with-figure</id>
    <title>Finding: 1 of 61 AI training data lawsuits tracked has a disclosed settlement figure.</title>
    <link href="https://data.blomega.com/findings#litigation-settled-with-figure"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>1 of 61 AI training data lawsuits tracked has a disclosed settlement figure. Method: Settlement value is kept separate from damages sought and damages awarded. A figure pleaded in a complaint is not counted.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#vendors-price-opacity</id>
    <title>Finding: Only 1 of 49 AI training data vendors tracked publishes what they charge buyers (Prolific); 4 publish a figure for what they pay contributors.</title>
    <link href="https://data.blomega.com/findings#vendors-price-opacity"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>Only 1 of 49 AI training data vendors tracked publishes what they charge buyers (Prolific); 4 publish a figure for what they pay contributors. Method: Buyer pricing counts as published only as a figure or a stated fee percentage on the vendor's own page. Quote-only, tiers without prices and prices charged to rights holders rather than AI buyers do not count. A row that only points to another vendor's figure is not counted twice.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#rates-two-markets</id>
    <title>Finding: Published hourly pay for AI training data work splits into two markets: crowd and annotation work at $3 to $20 per hour, and robot demonstration and teleoperation at $25 to $55 per</title>
    <link href="https://data.blomega.com/findings#rates-two-markets"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>Published hourly pay for AI training data work splits into two markets: crowd and annotation work at $3 to $20 per hour, and robot demonstration and teleoperation at $25 to $55 per hour. Method: Uses only rates published per hour in USD as a figure or range. Ranges report the full published bounds, not midpoints. Caps, floors, per-task rates and other currencies are excluded rather than converted.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#orgs-top-licensee</id>
    <title>Finding: OpenAI is the licensee in 22 of 75 tracked AI content licensing deals, more than any other company in the tracked set.</title>
    <link href="https://data.blomega.com/findings#orgs-top-licensee"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>OpenAI is the licensee in 22 of 75 tracked AI content licensing deals, more than any other company in the tracked set. Method: Counts deals where the organisation is the party receiving the licence. Organisation names are reconciled across spellings and legal suffixes.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#orgs-top-defendant</id>
    <title>Finding: OpenAI is a defendant in 13 of 61 tracked AI training data lawsuits, more than any other company in the tracked set. Public trackers list roughly 200 such cases, so this ranks the </title>
    <link href="https://data.blomega.com/findings#orgs-top-defendant"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>OpenAI is a defendant in 13 of 61 tracked AI training data lawsuits, more than any other company in the tracked set. Public trackers list roughly 200 such cases, so this ranks the tracked subset, not the whole docket. Method: Counts cases naming the organisation as a defendant. Consolidated cases are counted as the registry records them.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#models-licensors-unnamed</id>
    <title>Finding: 54 of 56 tracked EU training-data summaries that answer yes or other on licensed data name no licensor. 10 of those developers hold 48 announced licensing deals in this registry.</title>
    <link href="https://data.blomega.com/findings#models-licensors-unnamed"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>54 of 56 tracked EU training-data summaries that answer yes or other on licensed data name no licensor. 10 of those developers hold 48 announced licensing deals in this registry. Method: From the public summaries of training content that general-purpose AI model providers publish under EU AI Act Article 53(1)(d). 'Other' is counted with yes because those summaries describe partnerships in prose. Deals are joined to the developer organisation, not to a model: an announcement does not say which model the data trained.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#benchmarks-contamination-independent</id>
    <title>Finding: 12 of 38 tracked AI benchmarks have at least one contamination claim made by a party that is neither the benchmark's maintainer nor the model's developer. 10 of 38 have data establ</title>
    <link href="https://data.blomega.com/findings#benchmarks-contamination-independent"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>12 of 38 tracked AI benchmarks have at least one contamination claim made by a party that is neither the benchmark's maintainer nor the model's developer. 10 of 38 have data established as usable commercially. Method: A claim counts as independent only when the claimant is neither the benchmark maintainer nor the developer of the model named. A model developer's own disclosure of contamination is recorded but not counted. Commercial use is left unestablished when a licence covers only the compilation of material taken from elsewhere.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#listings-price-disclosure</id>
    <title>Finding: 37 of 79 AI training-data listings opened on data marketplaces showed a public price, and 18 tied that price to a volume of data. 0 of 25 listings on Defined.ai and Wirestock showe</title>
    <link href="https://data.blomega.com/findings#listings-price-disclosure"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>37 of 79 AI training-data listings opened on data marketplaces showed a public price, and 18 tied that price to a volume of data. 0 of 25 listings on Defined.ai and Wirestock showed a price. Method: Every listing card or dataset page actually opened on one day was counted. Priced means a currency amount visible without a login or a quote request. Category pages are sorted by the marketplace, so this is not a random sample.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#enforcement-priced</id>
    <title>Finding: 18 of 33 regulator actions over AI training data carry a money figure. The rest are orders, bans, reprimands and undertakings, where the regulator either has no fining power or did</title>
    <link href="https://data.blomega.com/findings#enforcement-priced"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>18 of 33 regulator actions over AI training data carry a money figure. The rest are orders, bans, reprimands and undertakings, where the regulator either has no fining power or did not use it. Method: Only actions whose subject is AI training data, model training on personal or copyrighted material, scraping for AI, or biometric data used for AI. A GDPR fine for a breach is out of scope. Amounts are as the regulator published them and are never converted between currencies, so the figures below are grouped by currency rather than summed. An amount that an appeal has erased is not counted as a price.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#models-crawler-named</id>
    <title>Finding: 39 of 76 EU public training-data summaries name the crawler that collected the web data. 23 name only a third-party corpus such as Common Crawl, and 14 name neither. A crawler nobo</title>
    <link href="https://data.blomega.com/findings#models-crawler-named"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>39 of 76 EU public training-data summaries name the crawler that collected the web data. 23 name only a third-party corpus such as Common Crawl, and 14 name neither. A crawler nobody names is a crawler nobody can block. Method: Each summary was read for a user-agent string, never inferred from the developer: OpenAI publishing a summary is not the summary naming GPTBot. A corpus named in the crawler field (Common Crawl, RefinedWeb, FineWeb) is counted as corpus, not as a crawler. Summaries that tick 'crawlers used: yes' and then answer the name field with 'NA', a hyperlink or a confidentiality clause are counted as not naming one.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#procurement-unit-price</id>
    <title>Finding: 0 of 40 public contract awards for AI data work state a price per unit of data. 38 state a total amount at all. An award notice is the one public filing that gives a buyer's price,</title>
    <link href="https://data.blomega.com/findings#procurement-unit-price"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>0 of 40 public contract awards for AI data work state a price per unit of data. 38 state a total amount at all. An award notice is the one public filing that gives a buyer's price, and it gives a contract total, not a rate per image, hour or word. Method: Awards were read from USAspending, UK Contracts Finder and EU TED, keeping only notices whose OBJECT is data collection, annotation or licensing for AI. Amounts are as filed, never converted between currencies, and obligated amounts are not mixed with ceilings in the medians. Two notices with image counts but no price were not divided out. The largest known contracts of this kind, such as the NGA SEQUOIA award, never reach these systems, so this is not the top of the market.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#vendor-pay-advertised-vs-reported</id>
    <title>Finding: 10 of 49 tracked data vendors now carry a worker-reported pay figure alongside what the vendor publishes, and 10 carry both. The two are listed side by side rather than compared, b</title>
    <link href="https://data.blomega.com/findings#vendor-pay-advertised-vs-reported"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>10 of 49 tracked data vendors now carry a worker-reported pay figure alongside what the vendor publishes, and 10 carry both. The two are listed side by side rather than compared, because they are quoted on different bases. Method: Reported pay comes from news investigations with documents or named sample sizes, union statements, audited surveys (Fairwork, ILO) and salary sites that state how many submissions a figure rests on. It is carried verbatim, never parsed into a number, because the quotes mix hourly, annual and per-task bases. Vendor pay is what the vendor publishes. Neither figure is adjusted for region or year.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#listings-price-per-hour</id>
    <title>Finding: Published prices per hour of training data: audio in EUR: median 214 per hour across 30 listings, from 35 to 3,027; video in GBP: median 60 per hour across 9 listings, from 6 to 2,</title>
    <link href="https://data.blomega.com/findings#listings-price-per-hour"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>Published prices per hour of training data: audio in EUR: median 214 per hour across 30 listings, from 35 to 3,027; video in GBP: median 60 per hour across 9 listings, from 6 to 2,886; audio in USD: median 25 per hour across 9 listings, from 25 to 250. Method: Only listings whose own page ties the price to a number of hours. Prices are the seller's list price in the seller's currency, not converted and not negotiated. Catalogues that publish prices (ELRA, Opendatabay) dominate the sample, so this is the price of openly listed data, not of quote-only premium collections.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#benchmarks-canary-strings</id>
    <title>Finding: 7 of 36 AI benchmarks whose released files we scanned carry a canary string, the marker that makes contamination provable instead of arguable.</title>
    <link href="https://data.blomega.com/findings#benchmarks-canary-strings"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>7 of 36 AI benchmarks whose released files we scanned carry a canary string, the marker that makes contamination provable instead of arguable. Method: Each row was scanned in its released data files (Hugging Face parquet and rows API, repository tarballs, published artifacts), and the file and method are recorded on the row. Benchmarks whose files were not scanned are excluded from both counts. Every canary found is the same BIG-bench GUID, reused by later benchmarks.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#datasets-license-chain-conflict</id>
    <title>Finding: 8 of 27 datasets whose rights chain was read back to the material they are built from carry terms that conflict: a permissive licence at the top over sources that restrict commerci</title>
    <link href="https://data.blomega.com/findings#datasets-license-chain-conflict"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>8 of 27 datasets whose rights chain was read back to the material they are built from carry terms that conflict: a permissive licence at the top over sources that restrict commercial use, or a card that grants and withholds it in the same page. Method: Each row was checked against the LICENSE file, the dataset card and the terms of the named upstream sources, not against a repository tag. Both sides of a conflict are recorded and no verdict is picked, so commercial use stays unestablished for these rows.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#crawl-blocks-licensed-still-blocking</id>
    <title>Finding: 13 of 65 sites that signed a licensing deal in this registry still disallow the crawler of a company they licensed to. A licence is a feed, not a crawl permit.</title>
    <link href="https://data.blomega.com/findings#crawl-blocks-licensed-still-blocking"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>13 of 65 sites that signed a licensing deal in this registry still disallow the crawler of a company they licensed to. A licence is a feed, not a crawl permit. Method: Joins the measured robots.txt state to this registry's licensing deals by hostname and licensee. Counted only when the blocked operator is one the site actually signed with. Says nothing about what a contract permits privately: a deal can deliver data by feed or dump while the public crawler stays blocked. Not a first: FT Strategies published the same join on 2026-07-14 across 70 publishers and named AP, the FT, El Pais and USA Today. This names 13 sites, per crawler and per deal, and is a dated state that can be diffed against theirs.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#orgs-license-and-sued</id>
    <title>Finding: 12 companies in the tracked set both license content for AI and are defendants in individual AI training data lawsuits: Adobe, Amazon, Apple, ElevenLabs, Google, Meta, Microsoft, N</title>
    <link href="https://data.blomega.com/findings#orgs-license-and-sued"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>12 companies in the tracked set both license content for AI and are defendants in individual AI training data lawsuits: Adobe, Amazon, Apple, ElevenLabs, Google, Meta, Microsoft, NVIDIA, OpenAI, Perplexity AI, Suno, Udio. Method: A company counts when it is the licensee in at least one tracked content deal AND a named defendant in at least one individually recorded case. A bundled record covering several suits from one source is excluded.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#orgs-sued-and-licensed</id>
    <title>Finding: 5 rights holders in the tracked set have both sued an AI company for copyright infringement and entered a content licensing deal with an AI company: Disney, Getty Images, New York </title>
    <link href="https://data.blomega.com/findings#orgs-sued-and-licensed"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>5 rights holders in the tracked set have both sued an AI company for copyright infringement and entered a content licensing deal with an AI company: Disney, Getty Images, New York Times, Reddit, UMG Recordings. UMG Recordings licensed to Udio, a company it sued. Method: A rights holder counts when it is the named plaintiff in a recorded case AND the licensor in a recorded deal. Deals that were later terminated or that exclude training are still counted as deals, and are listed as caveats.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/findings#registry-corroborated</id>
    <title>Finding: Of 1148 records served, 26 are sourced only to Blomega properties and are excluded from every market statistic here. Of the remaining 1122: 287 carry a second independent source, a</title>
    <link href="https://data.blomega.com/findings#registry-corroborated"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="finding"/>
    <summary>Of 1148 records served, 26 are sourced only to Blomega properties and are excluded from every market statistic here. Of the remaining 1122: 287 carry a second independent source, and 523 are primary measurements of a single legitimate publisher, where a second party cannot exist by construction. Method: Independent means a different party, not merely a different domain: a paper's own project page or an AI summary of the paper does not count. A site's own robots.txt, a marketplace's own listing, a developer's own EU summary and a state's own broker registry have exactly one legitimate publisher, so those rows carry a recorded method, URL and date instead of a corroborating source. Reading the ratio as 'the rest is unverified' double-counts that distinction.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/sensors#sensor:gelslim-30</id>
    <title>sensor: GelSlim 3.0</title>
    <link href="https://data.blomega.com/registry/sensors#sensor:gelslim-30"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="sensor"/>
    <summary>maker: Massachusetts Institute of Technology; type: vision-based-tactile; price usd: $25 per-unit. Update: arXiv 2103.12269 Table I, 'Comparison of GelSlim 3.0, GelSlim 2.0, GelSight, Digit and Omnitact', row 'Cost Components [$]' gives GelSlim 3.0 = 25*, with the table footnote '(*Considering the manufacturing of 1000 pieces)'. The same table gives GelSight 30 and Digit 15*, and Omnitact 600. The unit b. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/sensors#sensor:gelsight-mini</id>
    <title>sensor: GelSight Mini</title>
    <link href="https://data.blomega.com/registry/sensors#sensor:gelsight-mini"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="sensor"/>
    <summary>maker: GelSight Inc.; type: vision-based-tactile; price usd: $510. Update: Price re-read on the maker's own store 2026-09-18: $510.00. The $499 figure is the launch price and is corroborated for that date by an independent table (arXiv 2602.00514, Liu et al.).. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/sensors#sensor:forceband</id>
    <title>sensor: ForceBand</title>
    <link href="https://data.blomega.com/registry/sensors#sensor:forceband"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="sensor"/>
    <summary>maker: Amazon FAR, University of Maryland, Johns Hopkins; type: sEMG wristband, force inferred not measured; price usd: $300. Update: arXiv 2606.26093 states: 'The mechanical structure can be fabricated with common tools such as a commercial 3D printer, while the electronics are modular and sourced from readily available parts. The total cost can be as low as $300 depending on supplier. We open-source the complete bill of material. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/sensors#sensor:bota-rokubi</id>
    <title>sensor: Bota Rokubi</title>
    <link href="https://data.blomega.com/registry/sensors#sensor:bota-rokubi"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="sensor"/>
    <summary>maker: Bota Systems; type: force-torque; price usd: $5,408. Update: The maker publishes a price after all, just not on the page the registry cites. botasys.com/force-torque-sensors/rokubi says only 'Get a Quote', but Bota runs a storefront at shop.botasys.com that lists Rokubi Gen A with an Add to Cart button and a figure. Currency is Swiss francs; no USD figure is . Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/sensors#sensor:9dtact</id>
    <title>sensor: 9DTact</title>
    <link href="https://data.blomega.com/registry/sensors#sensor:9dtact"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="sensor"/>
    <summary>maker: Lin, Zhang et al.; type: vision-based-tactile; price usd: $15 per-unit. Update: 9DTact is an open-hardware sensor with no vendor, so there is no quote to gate. The designers publish the cost themselves: arXiv 2308.14277 Table I ('Comparison of GelSight, GelSlim 3.0, DIGIT, GelSight-Mini, DTact, and 9DTact') gives the row '9DTact (Ours) ... Cost[$] 15'. The table's own footnote . Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/rates#rate:transcription-rev</id>
    <title>rate: transcription</title>
    <link href="https://data.blomega.com/registry/rates#rate:transcription-rev"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="rate"/>
    <summary>platform: Rev; rate usd: $1.99; region: US. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/rates#rate:rlhf-expert-mercor-entrymidexpert</id>
    <title>rate: rlhf-expert</title>
    <link href="https://data.blomega.com/registry/rates#rate:rlhf-expert-mercor-entrymidexpert"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="rate"/>
    <summary>platform: Mercor; rate usd: $75 to $200; region: Global. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/rates#rate:image-annotation-amazon-mechanical-turk</id>
    <title>rate: image-annotation</title>
    <link href="https://data.blomega.com/registry/rates#rate:image-annotation-amazon-mechanical-turk"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="rate"/>
    <summary>platform: Amazon Mechanical Turk; rate usd: $2.83 per-hour; region: Global. Update: Amazon Mechanical Turk permanently closes 2026-09-30 (its own site, checked 2026-09-17). Requesters can approve or reject until 2026-10-30, after which work auto-approves; transaction history stays available until 2027-01-28. This rate is now historical.. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:smollm3-3b</id>
    <title>model: SmolLM3-3B</title>
    <link href="https://data.blomega.com/registry/models#model:smollm3-3b"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Hugging Face; release date: 2025-07-08; weights: open; licensed data declared: no; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? No'. 3.1 names the corpus rather than a crawler: 'All crawl-based data in the datasets uses the CommonCrawl archives which comply with robots.txt. Some datasets such as the Stack v2 additionally . Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:pllum-2512-instruct-4b-8b-12b-70b</id>
    <title>model: PLLuM 2512 instruct (4B, 8B, 12B, 70B)</title>
    <link href="https://data.blomega.com/registry/models#model:pllum-2512-instruct-4b-8b-12b-70b"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Ministry of Digital Affairs of Poland; release date: 2026-05-21; weights: open; licensed data declared: yes; licensed data named: False; crawler named: CCBot. Update: 2.3 'Dane zebrane w procesach automatycznego przeszukania i pozyskania z internetu': 'Czy przez dostawce lub w jego imieniu byly wykorzystywane roboty indeksujace? X tak'. The name field, 'Jezeli tak, nalezy podac nazwe (nazwy) robota indeksujacego/jego identyfikator', answers: 'Uzyto dedykowanych r. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:pllum-2512-base-4b-8b-12b</id>
    <title>model: PLLuM 2512 base (4B, 8B, 12B)</title>
    <link href="https://data.blomega.com/registry/models#model:pllum-2512-base-4b-8b-12b"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Ministry of Digital Affairs of Poland; release date: 2026-05-21; weights: open; licensed data declared: yes; licensed data named: False; crawler named: CCBot. Update: 2.3 'Dane zebrane w procesach automatycznego przeszukania i pozyskania z internetu': 'Czy przez dostawce lub w jego imieniu byly wykorzystywane roboty indeksujace? X tak'. The name field, 'Jezeli tak, nalezy podac nazwe (nazwy) robota indeksujacego/jego identyfikator', answers: 'Uzyto dedykowanych r. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:pleias-1-0-350m-1-2b-3b-preview</id>
    <title>model: Pleias 1.0 (350m, 1.2b, 3b Preview)</title>
    <link href="https://data.blomega.com/registry/models#model:pleias-1-0-350m-1-2b-3b-preview"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Pleias; release date: 2024-12; weights: open; licensed data declared: no; licensed data named: False; crawler named: EDGAR-Crawler. Update: 2.3: 'Were crawlers used by the provider or on behalf of? [x] Yes'. The name field answers in full: 'PLEIAS does not operate a general-purpose web crawler and has not crawled the open web at large. No crawler product or user-agent identifier is published by PLEIAS.' Content was obtained by '(i) inte. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-4-reasoning</id>
    <title>model: Phi-4-reasoning</title>
    <link href="https://data.blomega.com/registry/models#model:phi-4-reasoning"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2025-04-30; weights: open; licensed data declared: yes; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. Sources are 'Prompts sourced from publicly available websites, existing datasets, and licensed collections. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-4-multimodal-instruct</id>
    <title>model: Phi-4-multimodal-instruct</title>
    <link href="https://data.blomega.com/registry/models#model:phi-4-multimodal-instruct"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2025-02; weights: open; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. Sources are given as 'Anonymized in-house speech-text pairs with strong and weak transcriptions, selected . Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-4-mini-instruct-reasoning-flash-reasoning</id>
    <title>model: Phi-4-mini (instruct, reasoning, flash-reasoning)</title>
    <link href="https://data.blomega.com/registry/models#model:phi-4-mini-instruct-reasoning-flash-reasoning"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2025-04-29; weights: open; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. The card describes only the reasoning variant's data: 'exclusively of synthetic mathematical content gener. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-4</id>
    <title>model: Phi-4</title>
    <link href="https://data.blomega.com/registry/models#model:phi-4"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2024-12-12; weights: open; licensed data declared: yes; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. 1.3.1.B lists source categories only ('Publicly available documents filtered rigorously for quality', 'Acq. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-3-vision-and-phi-3-5-vision-instruct</id>
    <title>model: Phi-3 Vision and Phi-3.5 Vision Instruct</title>
    <link href="https://data.blomega.com/registry/models#model:phi-3-vision-and-phi-3-5-vision-instruct"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2024-05-21; weights: open; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. 1.3.1.D gives only 'Selected image-text interleaved data and newly created image data including charts, ta. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-3-mini-small-medium-instruct</id>
    <title>model: Phi-3 (mini, small, medium instruct)</title>
    <link href="https://data.blomega.com/registry/models#model:phi-3-mini-small-medium-instruct"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2024-05-21; weights: open; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. 1.3.1.B gives only 'Our training data includes a wide variety of sources and is a combination of publicly . Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-2</id>
    <title>model: Phi-2</title>
    <link href="https://data.blomega.com/registry/models#model:phi-2"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2023-12-12; weights: open; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. 1.3.1.B names 'NLP synthetic data created by Azure OpenAI GPT-3.5 and filtered web data from Falcon Refine. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-1-5</id>
    <title>model: Phi-1.5</title>
    <link href="https://data.blomega.com/registry/models#model:phi-1-5"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2023-10-09; weights: open; licensed data declared: no; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. 1.3.1.B says 'Same data sources as phi-1, augmented with various NLP synthetic texts generated by gpt-3.5-. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:phi-1</id>
    <title>model: Phi-1</title>
    <link href="https://data.blomega.com/registry/models#model:phi-1"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2023-09-10; weights: open; licensed data declared: no; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. The card has no crawled-data section at all: its 2.3 is 'Personal Information', and the words crawler, crawl and scrape do not appear anywhere in it. 1.3.1.B names the sources instead: 'subsets of Python codes from The Stack v1.2, Q&amp;A content from StackOve. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:muse-spark</id>
    <title>model: Muse Spark</title>
    <link href="https://data.blomega.com/registry/models#model:muse-spark"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Meta; release date: 2026-04-08; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? Yes'. 'If yes, specify crawler name(s)/identifier(s): It is Meta's policy to provide information on Meta's web crawlers through Meta's developer center (https://developers.facebook.com/docs/shari. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:muse-image</id>
    <title>model: Muse Image</title>
    <link href="https://data.blomega.com/registry/models#model:muse-image"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Meta; release date: 2026-07-07; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? Yes'. 'If yes, specify crawler name(s)/identifier(s): It is Meta's policy to provide information on Meta's web crawlers through Meta's developer center (https://developers.facebook.com/docs/shari. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:muse-glimmer</id>
    <title>model: Muse Glimmer</title>
    <link href="https://data.blomega.com/registry/models#model:muse-glimmer"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Meta; release date: 2026-08-10; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? Yes'. 'If yes, specify crawler name(s)/identifier(s): It is Meta's policy to provide information on Meta's web crawlers through Meta's developer center (https://developers.facebook.com/docs/shari. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:mistral-small-4</id>
    <title>model: Mistral Small 4</title>
    <link href="https://data.blomega.com/registry/models#model:mistral-small-4"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Mistral AI; release date: 2026-03-16; weights: open; licensed data declared: other; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): NA.' Purposes: 'Crawlers were used to collect publicly available sources on the internet.' Behaviour: 'Our crawlers are designed to respe. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:mistral-large-3</id>
    <title>model: Mistral Large 3</title>
    <link href="https://data.blomega.com/registry/models#model:mistral-large-3"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Mistral AI; release date: 2025-12-02; weights: open; licensed data declared: other; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): NA.' Purposes: 'Crawlers were used to collect publicly available sources on the internet.' Behaviour: 'Our crawlers are designed to respe. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:ministral-3-8b-base-instruct-reasoning</id>
    <title>model: Ministral 3 8B (Base, Instruct, Reasoning)</title>
    <link href="https://data.blomega.com/registry/models#model:ministral-3-8b-base-instruct-reasoning"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Mistral AI; release date: 2025-12-02; weights: open; licensed data declared: other; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): NA.' Purposes: 'Crawlers were used to collect publicly available sources on the internet.' Behaviour: 'Our crawlers are designed to respe. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:ministral-3-3b-base-instruct-reasoning</id>
    <title>model: Ministral 3 3B (Base, Instruct, Reasoning)</title>
    <link href="https://data.blomega.com/registry/models#model:ministral-3-3b-base-instruct-reasoning"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Mistral AI; release date: 2025-12-02; weights: open; licensed data declared: other; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): NA.' Purposes: 'Crawlers were used to collect publicly available sources on the internet.' Behaviour: 'Our crawlers are designed to respe. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:ministral-3-14b-base-instruct-reasoning</id>
    <title>model: Ministral 3 14B (Base, Instruct, Reasoning)</title>
    <link href="https://data.blomega.com/registry/models#model:ministral-3-14b-base-instruct-reasoning"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Mistral AI; release date: 2025-12-02; weights: open; licensed data declared: other; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): NA.' Purposes: 'Crawlers were used to collect publicly available sources on the internet.' Behaviour: 'Our crawlers are designed to respe. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:midjourney-image-and-video-family-v8-v8-1-v8-2</id>
    <title>model: Midjourney Image and Video family (V8, V8.1, V8.2)</title>
    <link href="https://data.blomega.com/registry/models#model:midjourney-image-and-video-family-v8-v8-1-v8-2"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Midjourney; release date: 2026-03-17; weights: closed; licensed data declared: yes; licensed data named: False. Update: 2.3: 'Were crawlers used by the provider or on behalf of? [x] Yes'. The field 'If yes, specify crawler name(s)/identifier(s):' is answered 'Deals are bound by confidentiality obligations', a sentence about licensing deals placed in the crawler-name box. Purposes: 'Crawlers are used to obtain publicl. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:mai-ds-r1</id>
    <title>model: MAI-DS-R1</title>
    <link href="https://data.blomega.com/registry/models#model:mai-ds-r1"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2025-04; weights: open; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. No crawled-data section (its 2.3 is 'Personal Information'); crawler, crawl and scrape do not appear. The card covers post-training only: '110k Safety and Non-Compliance examples from the Tulu 3 SFT dataset and ~350k multilingual examples internally develo. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:magma-8b</id>
    <title>model: Magma-8B</title>
    <link href="https://data.blomega.com/registry/models#model:magma-8b"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2025-02-19; weights: open; licensed data declared: no; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. No crawled-data section exists in the card (its 2.3 is 'Personal Information') and the words crawler, crawl and scrape do not appear. Sources are named as datasets in 1.3.1: Open-X-Embodiment, Ego4D, Epic-Kitchens, Something-Something v2, ShareGPT4V, SeeCl. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:luciole-instruct-1-1-1b-8b-23b</id>
    <title>model: Luciole Instruct 1.1 (1B, 8B, 23B)</title>
    <link href="https://data.blomega.com/registry/models#model:luciole-instruct-1-1-1b-8b-23b"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: LINAGORA; release date: 2026-07-09; weights: open; licensed data declared: no; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] No'. Unlike the Luciole Base summary this one contains no crawler name anywhere: the strings CCBot, Common Crawl and robots.txt do not appear. Post-training content is described as 'pre-packa. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:luciole-base-1b-8b-23b</id>
    <title>model: Luciole Base (1B, 8B, 23B)</title>
    <link href="https://data.blomega.com/registry/models#model:luciole-base-1b-8b-23b"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: LINAGORA; release date: 2026-06-02; weights: open; licensed data declared: no; licensed data named: False; crawler named: CCBot. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] No'. LINAGORA operated no crawler. The crawler name appears in 3.1, describing the retroactive opt-out filter applied to every web-derived dataset: 'Robots.txt files were retrieved from the C. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:inkling-small</id>
    <title>model: Inkling-Small</title>
    <link href="https://data.blomega.com/registry/models#model:inkling-small"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Thinking Machines Lab; release date: 2026-07-30; weights: open; licensed data declared: yes; licensed data named: False. Update: 2.3: 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): N/A'. The Inkling-Small text is identical to the Inkling summary apart from the model name, including the crawl period 'From 2025 to 2026'. 2.1 names 'Common Crawl (https://commoncrawl. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:inkling</id>
    <title>model: Inkling</title>
    <link href="https://data.blomega.com/registry/models#model:inkling"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Thinking Machines Lab; release date: 2026-07-15; weights: open; licensed data declared: yes; licensed data named: False. Update: 2.3: 'Were crawlers used by the provider or on behalf of? [x] Yes'. 'If yes, specify crawler name(s)/identifier(s): N/A'. Purposes: 'Crawlers were used to download content from publicly available sources from the internet for the purpose of model training'. Behaviour: 'Thinking Machines Lab's policy. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:grin-moe-and-phi-moe-tiny-mini-instruct</id>
    <title>model: GRIN-MoE and Phi MoE (tiny, mini) instruct</title>
    <link href="https://data.blomega.com/registry/models#model:grin-moe-and-phi-moe-tiny-mini-instruct"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Microsoft; release date: 2024-09-18; weights: open; licensed data declared: no; licensed data named: False. Update: Microsoft 'Data Summary' card, three pages. Its section 2.3 is 'Personal Information', not crawled data: the card carries no crawled-data question at all, and the words crawler, crawl and scrape do not appear. Sources are given only as '1.3.1.B Text training data content: Our training data includes . Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:gemma-4-family</id>
    <title>model: Gemma 4 (family)</title>
    <link href="https://data.blomega.com/registry/models#model:gemma-4-family"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Google; release date: 2026-04; weights: open; licensed data declared: yes; licensed data named: False; crawler named: Google-Extended. Update: 2.3: crawlers Yes; 'If yes, specify crawler name(s)/identifier(s): See list of crawlers here.' The Gemma 4 text is word-for-word the Gemini 3 Pro text. The only crawler string printed is in 3.1: 'our Google-Extended control lets web publishers manage whether content Google crawls from their sites ma. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:gemini-3-pro-family</id>
    <title>model: Gemini 3 Pro (family)</title>
    <link href="https://data.blomega.com/registry/models#model:gemini-3-pro-family"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Google; release date: 2025-11; weights: closed; licensed data declared: yes; licensed data named: False; crawler named: Google-Extended. Update: 2.3: 'Were crawlers used by the provider or on behalf of? Yes'. The field 'If yes, specify crawler name(s)/identifier(s)' is answered with a bare pointer, 'See list of crawlers here.', and the domains field repeats 'See list of our common crawlers here. Google's common crawlers obey robots.txt rules. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:flux-3</id>
    <title>model: FLUX 3</title>
    <link href="https://data.blomega.com/registry/models#model:flux-3"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Black Forest Labs; release date: 2026-07-16; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] No'. No crawler sub-fields are printed. The only 'common crawl' strings in the document are the template's own boilerplate ('platforms such as common crawl that are covered under Section 2.1'. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:fibo-fibo-fibo-lite-fibo-edit</id>
    <title>model: FIBO (FIBO, FIBO Lite, FIBO Edit)</title>
    <link href="https://data.blomega.com/registry/models#model:fibo-fibo-fibo-lite-fibo-edit"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Bria AI; release date: 2025-10-29; weights: open; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on their behalf? No. Bria does not engage in web crawling or scraping. Accordingly no crawler identifiers, collection period or domain name list falls to be disclosed.' The annex 'No Web-Crawling Policy' repea. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:fastwebmiia-fastwebmiia-7b-fastwebmiia-7b-2603</id>
    <title>model: FastwebMIIA (FastwebMIIA-7B, FastwebMIIA-7B-2603)</title>
    <link href="https://data.blomega.com/registry/models#model:fastwebmiia-fastwebmiia-7b-fastwebmiia-7b-2603"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Fastweb; release date: 2025-05-29; weights: open; licensed data declared: yes; licensed data named: True. Update: Row U of the Italian narrative adaptation, 'Identificazione dei crawler, loro scopo e comportamento', is answered with three dataset names, not crawlers: 'I principali Crawler utilizzati sono: 1) Common Crawl 2) Red Pajama 3) FineWeb Edu'. Row W repeats the same three names as the most relevant inte. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:elevenlabs-model-families-tts-voice-stt-music-non-speech</id>
    <title>model: ElevenLabs model families (TTS, Voice, STT, Music, Non-Speech)</title>
    <link href="https://data.blomega.com/registry/models#model:elevenlabs-model-families-tts-voice-stt-music-non-speech"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: ElevenLabs; weights: closed; licensed data declared: yes; licensed data named: False. Update: This is a California AB 2013 'Training Data Transparency Disclosure' (effective 1 January 2026), not the EU Article 53 template, and it has no crawled-data section at all. The words crawler, crawl and scrape do not appear in the document. Sources are described only as categories: 'Licensed voice rec. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:domyn-large</id>
    <title>model: Domyn Large</title>
    <link href="https://data.blomega.com/registry/models#model:domyn-large"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Domyn; release date: 2026-03-20; licensed data declared: no; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'We have not crawled, scraped, or otherwise directly compiled data from online sources ourselves or through third parties on our behalf'. 3.1 says only 'All open dataset have used web crawlers that honor machine-readable opt-out signals, such as ro. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:deepseek-v4-pro-and-flash</id>
    <title>model: DeepSeek-V4 (Pro and Flash)</title>
    <link href="https://data.blomega.com/registry/models#model:deepseek-v4-pro-and-flash"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: DeepSeek; release date: 2026-04-24; weights: open; licensed data declared: yes; licensed data named: False. Update: The crawled-data section of the template is missing from this summary: sections run 2.1 publicly available datasets, 2.2 private datasets, 2.3 User data, 2.4 Synthetic data, 2.5 Other sources. The only crawler reference is in 3.1: 'the crawler is designed to respect robots.txt instructions and other. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:command-a-c4ai-command-a-plus</id>
    <title>model: Command A+ (C4AI Command A Plus)</title>
    <link href="https://data.blomega.com/registry/models#model:command-a-c4ai-command-a-plus"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Cohere; release date: 2026-05-20; weights: open; licensed data declared: other; licensed data named: False. Update: 2.3: 'Were crawlers used by the provider or on behalf of? Yes'. The field 'If yes, specify crawler name(s)/identifier(s):' is answered with a link, not a name: 'Cohere makes information about its web crawlers available at https://docs.cohere.com/docs/cohere-web-crawlers.' Purposes adds 'Prior to Aug. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:bria-3-2</id>
    <title>model: Bria 3.2</title>
    <link href="https://data.blomega.com/registry/models#model:bria-3-2"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Bria AI; release date: 2025-06-10; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data Crawled and Scraped from Online Sources': 'Q: Were crawlers used by the provider or on behalf of? A: No'. An annex headed 'No Web-Crawling Policy' adds 'Bria does not and will not engage in web-crawling activities or utilize publicly [available web data]'. No crawler is named and no third-. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:baguettotron-and-monad</id>
    <title>model: Baguettotron and Monad</title>
    <link href="https://data.blomega.com/registry/models#model:baguettotron-and-monad"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Pleias; release date: 2025-11-10; weights: open; licensed data declared: no; licensed data named: False. Update: 2.3: 'Were crawlers used by the provider or on behalf of? [x] No'. Crawler name field: 'Not applicable. No crawler was used by PLEIAS or on its behalf for the training of these models.' Behaviour field: seed material came from 'machine-readable dumps published by the Wikimedia Foundation through its. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:apertus-v1-8b-70b-instruct</id>
    <title>model: Apertus v1 (8B, 70B, Instruct)</title>
    <link href="https://data.blomega.com/registry/models#model:apertus-v1-8b-70b-instruct"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Swiss AI Initiative; release date: 2025-09-02; weights: open; licensed data declared: no; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] No'. No crawler sub-fields follow. 3.1 refers to websites that opted out 'by specifying at least one of the common AI crawlers, at the time of January 2025' without naming one.. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:apertus-v1-5-8b-70b-instruct</id>
    <title>model: Apertus v1.5 (8B, 70B, Instruct)</title>
    <link href="https://data.blomega.com/registry/models#model:apertus-v1-5-8b-70b-instruct"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Swiss AI Initiative; release date: 2026-07-14; weights: open; licensed data declared: no; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? [x] Yes'. Every sub-field of 2.3 (crawler name/identifier, purposes, behaviour, period, domains) is absent from the document: the template jumps straight from the Yes tick to 2.4 User data. 3.1 s. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/models#model:adobe-firefly-image-model-5</id>
    <title>model: Adobe Firefly Image Model 5</title>
    <link href="https://data.blomega.com/registry/models#model:adobe-firefly-image-model-5"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="model"/>
    <summary>developer: Adobe; release date: 2025-10-28; weights: closed; licensed data declared: yes; licensed data named: False. Update: 2.3 'Data crawled and scraped from online sources': 'Were crawlers used by the provider or on behalf of? No'. Purposes and behaviour both 'N/A'. The content field says 'Adobe did not crawl online sources. Instead, we searched specifically for content licensed under CC0 ... within select sites'. 'Sum. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:woulard-v-uncharted-labs</id>
    <title>litigation: Woulard v. Uncharted Labs</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:woulard-v-uncharted-labs"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: S.D.N.Y. (transferred from N.D. Ill.); defendant: Uncharted Labs; stage: motion to dismiss; filed: 2025-10-15. Update: Both dockets confirmed on CourtListener 2026-09-18: N.D. Ill. 1:25-cv-12613, filed 2025-10-15, terminated 2026-08-04, and S.D.N.Y. 1:26-cv-07968, filed 2026-09-12, not terminated. A sibling action, Woulard v. Suno, Inc., is N.D. Ill. 1:25-cv-12684 and is a different record. No settlement on either d. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:wikihow-v-openai</id>
    <title>litigation: wikiHow v. OpenAI</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:wikihow-v-openai"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: S.D.N.Y.; defendant: OpenAI Inc., eight affiliates; stage: filed; filed: 2026-08-21. Update: Docket confirmed on CourtListener 2026-09-18: S.D.N.Y. 1:26-cv-07171, wikiHow, Inc. v. OpenAI, Inc., filed 2026-08-21, Judge Sidney H. Stein, not terminated. No settlement on the docket.. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:umg-recordings-v-uncharted-labs-udio</id>
    <title>litigation: UMG Recordings v. Uncharted Labs (Udio)</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:umg-recordings-v-uncharted-labs-udio"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: S.D.N.Y.; defendant: Udio; stage: settled; filed: 2024-06-24. Update: Docket confirmed on CourtListener 2026-09-18: S.D.N.Y. 1:24-cv-04777, UMG Recordings, Inc. v. Uncharted Labs, Inc., filed 2024-06-24, Judge Alvin K. Hellerstein, not terminated. settlement_usd stays null: UMG's 2025-10-29 resolution was a licensing partnership rather than a disclosed cash payment, a. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:umg-recordings-v-suno</id>
    <title>litigation: UMG Recordings v. Suno</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:umg-recordings-v-suno"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: D. Mass.; defendant: Suno; stage: discovery; filed: 2024-06-24. Update: Docket confirmed on CourtListener 2026-09-18: D. Mass. 1:24-cv-11611-FDS, filed 2024-06-24, Judge F. Dennis Saylor IV, not terminated. settlement_usd stays null: the only settlement here is Warner's, its terms are sealed, and a magistrate has refused to disclose them to the remaining plaintiffs, so . Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:times-publishing-company-v-microsoft</id>
    <title>litigation: Times Publishing Company v. Microsoft</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:times-publishing-company-v-microsoft"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: S.D.N.Y.; defendant: Microsoft, OpenAI entities; stage: filed; filed: 2026-09-16. Update: Docket confirmed on CourtListener 2026-09-18: S.D.N.Y. 1:26-cv-08082, Times Publishing Company v. Microsoft Corporation, filed 2026-09-16, not terminated. chatgptiseatingtheworld counts it as the 143rd US AI copyright suit and expects it to be drawn into MDL 3143. No settlement on the docket.. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:the-intercept-v-openai</id>
    <title>litigation: The Intercept v. OpenAI</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:the-intercept-v-openai"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: S.D.N.Y.; defendant: OpenAI, Microsoft; stage: summary judgment. Update: Docket confirmed on CourtListener 2026-09-18: S.D.N.Y. 1:24-cv-01515, The Intercept Media, Inc. v. OpenAI, Inc., filed 2024-02-28, not terminated; a matching J.P.M.L. entry exists for the transfer into MDL 3143. The 2025-02-20 opinion already in the record stands. No settlement on the docket.. Grade B.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:sullivan-v-openai</id>
    <title>litigation: Sullivan v. OpenAI</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:sullivan-v-openai"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: S.D.N.Y.; defendant: OpenAI, Microsoft; stage: stayed; filed: 2026-08-14. Update: Docket confirmed on CourtListener 2026-09-18: S.D.N.Y. 1:26-cv-06966, captioned Sullivan v. OpenAI Foundation, filed 2026-08-14, Judge Sidney H. Stein, not terminated. The 2026-09-16 stay already in the record is corroborated on the docket (ECF 23, an order staying the case entered from MDL 3143 doc. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:sullivan-v-meta-platforms</id>
    <title>litigation: Sullivan v. Meta Platforms</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:sullivan-v-meta-platforms"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: N.D. Cal.; defendant: Meta Platforms, Mark Zuckerberg; stage: filed; filed: 2026-07-02. Update: Docket confirmed on CourtListener 2026-09-18: N.D. Cal. 3:26-cv-06793, Sullivan v. Meta Platforms, Inc, filed 2026-07-02, Judge Vince Girdhari Chhabria, not terminated. 76 docket entries in under three months. No settlement on the docket.. Grade A.</summary>
  </entry>
  <entry>
    <id>https://data.blomega.com/registry/litigation#litigation:sony-music-publishing-v-anthropic</id>
    <title>litigation: Sony Music Publishing v. Anthropic</title>
    <link href="https://data.blomega.com/registry/litigation#litigation:sony-music-publishing-v-anthropic"/>
    <updated>2026-09-18T00:00:00Z</updated>
    <category term="litigation"/>
    <summary>court: N.D. Cal.; defendant: Anthropic, Dario Amodei, Benjamin Mann; stage: filed; filed: 2026-08-28. Update: Docket confirmed on CourtListener 2026-09-18: N.D. Cal. 5:26-cv-09217, Sony Music Publishing (US) LLC v. Anthropic PBC, filed 2026-08-28, not terminated. No settlement on the docket.. Grade A.</summary>
  </entry>
</feed>
