Blomega Data Refinery / registry
AI training opt-out mechanisms and their legal force
41 records, built 2026-09-18. Machine-readable: JSON and JSON-LD.
| Name | Mechanism | Operator | Legal force | Jurisdiction | Does not cover | Grade | Source |
|---|---|---|---|---|---|---|---|
| Adobe no training on customer content (Firefly) | platform-toggle | Adobe | contractual | global | Adobe Stock contributors' submissions are used for training (compensated); this is a policy commitment, not a user toggle; content analys... | B | source |
| Amazonbot robots.txt token | robots-token | Amazon | voluntary | global | Amzn-User user-initiated fetches may not follow all robots.txt rules; data already collected | B | source |
| Applebot-Extended robots.txt token | robots-token | Apple | voluntary | global | Does not stop Applebot crawling or Spotlight/Siri/Safari search inclusion; Apple does not address removal of previously crawled data | B | source |
| Bytespider robots.txt token | robots-token | ByteDance | voluntary | global | No operator documentation read that commits to honouring robots.txt; data already collected | B | source |
| California DROP (Delete Act, SB 362) | regulatory-request | California Privacy Protection Agency | enforceable-law | US-CA | First-party data held by companies that are not data brokers (including most AI labs and platforms); not a copyright or training opt-out;... | B | source |
| CAWG training and data mining assertion (successor to C2PA do-not-train) | standard | Creator Assertions Working Group, used inside C2PA manifests | voluntary | global | C2PA removed its own training-mining assertion in spec v2.0, so older c2pa.* labels are obsolete; metadata is easily stripped; no model d... | B | source |
| CCBot robots.txt token | robots-token | Common Crawl Foundation | voluntary | global | Past crawl snapshots already published and copied by downstream users; page does not address retroactive removal; downstream AI developer... | B | source |
| ChatGPT-User robots.txt token | robots-token | OpenAI | voluntary | global | OpenAI itself says robots.txt may not apply because fetches are user-initiated; not a training control | B | source |
| Claude-SearchBot robots.txt token | robots-token | Anthropic | voluntary | global | Training (ClaudeBot); already-indexed content | B | source |
| Claude-User robots.txt token | robots-token | Anthropic | voluntary | global | Training (ClaudeBot) and search indexing (Claude-SearchBot) are separate tokens | B | source |
| ClaudeBot robots.txt token | robots-token | Anthropic | voluntary | global | Anthropic frames the signal as excluding the site's future materials from training; no statement on removing already-collected data; does... | B | source |
| Cloudflare AI Scrapers and Crawlers block | network-block | Cloudflare | contractual | global | Only sites behind Cloudflare; ML detection is imperfect; user-initiated agents and headless browsers can blur categories; does nothing fo... | B | source |
| Cloudflare Content Signals Policy (Content-Signal in robots.txt) | standard | Cloudflare | voluntary | global | Cloudflare says signals are preferences, not countermeasures, and some companies may ignore them; ai-input left unset by default; content... | B | source |
| Cloudflare Pay per crawl | network-block | Cloudflare | contractual | global | Private beta; only sites on Cloudflare; flat domain-wide price, no training vs inference distinction; unregistered crawlers are simply bl... | B | source |
| EU AI Act Article 53(1)(c)-(d) GPAI copyright policy and training summary | legal-reservation | European Union | enforceable-law | EU | Does not itself create a new opt-out; transitional rules for models placed before 2025-08-02 not verified; does not require removal from ... | B | source |
| EU GPAI Code of Practice, Copyright chapter | regulatory-request | European Commission AI Office | voluntary | EU | Non-signatories; voluntary instrument (a compliance route for Article 53, not the law itself); data already collected | B | source |
| EU TDM rights reservation (DSM Directive Article 4(3)) | legal-reservation | European Union (Directive (EU) 2019, 790) | enforceable-law | EU | Research organisations under Article 3 (no opt-out); works mined before the reservation was made; training performed outside the EU; uncl... | B | source |
| GDPR Article 21 right to object (general) | regulatory-request | European Union, mirrored in UK GDPR | enforceable-law | EU | Only personal data, not copyright in works; controller may continue if it shows compelling legitimate grounds; unlearning from trained mo... | B | source |
| Google search snippet controls (nosnippet, data-nosnippet, max-snippet, noindex) for AI Overviews and AI Mode | robots-token | voluntary | global | Cannot opt out of AI features while keeping full Search snippets; recrawl can take days to months; does not govern Gemini training (Googl... | B | source | |
| Google-CloudVertexBot robots.txt token | robots-token | voluntary | global | Not a general training opt-out; no effect on Search or other products | B | source | |
| Google-Extended robots.txt token | robots-token | voluntary | global | Does not crawl separately (reuses Googlebot); does NOT affect Google Search, AI Overviews or AI Mode, which are controlled only via Googl... | B | source | |
| GPTBot robots.txt token | robots-token | OpenAI | voluntary | global | Data already collected before the disallow; content reaching OpenAI via third-party datasets or licensed sources; does not affect ChatGPT... | B | source |
| IETF AI Preferences (aipref) vocabulary and attachment | standard | IETF aipref working group | proposed | global | Charter explicitly excludes technical enforcement; not yet an RFC; no retroactive effect | B | source |
| Known Agents (formerly Dark Visitors) robots.txt agent directory | network-block | Known Agents | voluntary | global | Covers crawler identity and robots.txt only; no legal force, platform toggles, regulatory routes, or data-already-collected gap | B | source |
| LinkedIn Data for Generative AI Improvement setting | platform-toggle | contractual | global | Does not affect training that already took place; feedback data and non-generative models need the separate Data Processing Objection form | B | source | |
| Meta AI training objection (GDPR Article 21 route) | regulatory-request | Meta Platforms | enforceable-law | EU | Private messages and under-18 data already excluded; announcement does not say objection removes data already used; content about you pos... | B | source |
| Meta-ExternalAgent robots.txt token | robots-token | Meta | voluntary | global | Meta's page does not explicitly state robots.txt compliance for this agent; Meta-ExternalFetcher (user-initiated) may bypass robots.txt; ... | B | source |
| Meta-ExternalFetcher token | robots-token | Meta | voluntary | global | Meta states it may bypass robots.txt | B | source |
| noai / noimageai directives (meta robots and X-Robots-Tag) | standard | DeviantArt | voluntary | global | No crawler operator documentation read commits to honouring it; DeviantArt acknowledges it cannot technically prevent scraping; content a... | B | source |
| noarchive meta tag as AI-training signal (Amazon) | robots-token | Amazon | voluntary | global | Only Amazon agents documented; page-level only | B | source |
| NOCACHE and NOARCHIVE meta tags for Bing Chat / Microsoft generative AI | robots-token | Microsoft | voluntary | global | Default (no tag) permits training use; announcement dates from 2023 and product names have since changed (Copilot); models already trained | B | source |
| OAI-SearchBot robots.txt token | robots-token | OpenAI | voluntary | global | Training (governed by GPTBot); user-initiated fetches (ChatGPT-User) | B | source |
| Perplexity-User token | robots-token | Perplexity | voluntary | global | Operator states it generally ignores robots.txt, so the token is identification only, not an opt-out | B | source |
| PerplexityBot robots.txt token | robots-token | Perplexity | voluntary | global | Perplexity-User fetches, which Perplexity says generally ignore robots.txt; undeclared crawling observed by Cloudflare | B | source |
| RSL (Really Simple Licensing) | standard | RSL Collective, RSL Internet Collective | voluntary | global | No AI developer commitment to honour it listed on the site; licensing enforcement depends on CDN partners or contracts | B | source |
| Spawning ai.txt and Do Not Train registry | standard | Spawning | voluntary | global | Could not verify current status: spawning.ai showed an under-maintenance page on 2026-09-16; honouring parties not verified | B | source |
| Squarespace Block known artificial intelligence crawlers | platform-toggle | Squarespace | voluntary | global | Only cooperative crawlers; no retroactive removal; no page-level control | B | source |
| TDM Reservation Protocol (TDMRep) | standard | W3C TDM Reservation Protocol Community Group | voluntary | EU | Is a W3C Community Group report, not a W3C Standard; no crawler operator documentation read commits to reading it; content mined before p... | B | source |
| UK TDM opt-out exception (proposed, abandoned) | legal-reservation | UK Government | proposed | UK | Not law: March 2026 report says a broad exception with opt-out is no longer the preferred way forward; current law (CDPA s29A) permits TD... | B | source |
| WordPress.com Prevent third-party sharing | platform-toggle | Automattic | contractual | global | Robots.txt part depends on AI platforms honouring it; per-site setting; no statement on data already shared | B | source |
| YouTube third-party training setting | platform-toggle | YouTube | contractual | global | Default off means no third-party permission, but it does not stop scraping by parties outside the programme; page does not cover Google's... | B | source |
Cite as: Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Published by
Blomega (Wikidata Q141048865). Unknown values are shown as
not established, never guessed.