Blomega Data Refinery / registry

AI training opt-out mechanisms and their legal force

41 records, built 2026-09-18. Machine-readable: JSON and JSON-LD.

NameMechanismOperatorLegal forceJurisdictionDoes not coverGradeSource
Adobe no training on customer content (Firefly)platform-toggleAdobecontractualglobalAdobe Stock contributors' submissions are used for training (compensated); this is a policy commitment, not a user toggle; content analys...Bsource
Amazonbot robots.txt tokenrobots-tokenAmazonvoluntaryglobalAmzn-User user-initiated fetches may not follow all robots.txt rules; data already collectedBsource
Applebot-Extended robots.txt tokenrobots-tokenApplevoluntaryglobalDoes not stop Applebot crawling or Spotlight/Siri/Safari search inclusion; Apple does not address removal of previously crawled dataBsource
Bytespider robots.txt tokenrobots-tokenByteDancevoluntaryglobalNo operator documentation read that commits to honouring robots.txt; data already collectedBsource
California DROP (Delete Act, SB 362)regulatory-requestCalifornia Privacy Protection Agencyenforceable-lawUS-CAFirst-party data held by companies that are not data brokers (including most AI labs and platforms); not a copyright or training opt-out;...Bsource
CAWG training and data mining assertion (successor to C2PA do-not-train)standardCreator Assertions Working Group, used inside C2PA manifestsvoluntaryglobalC2PA removed its own training-mining assertion in spec v2.0, so older c2pa.* labels are obsolete; metadata is easily stripped; no model d...Bsource
CCBot robots.txt tokenrobots-tokenCommon Crawl FoundationvoluntaryglobalPast crawl snapshots already published and copied by downstream users; page does not address retroactive removal; downstream AI developer...Bsource
ChatGPT-User robots.txt tokenrobots-tokenOpenAIvoluntaryglobalOpenAI itself says robots.txt may not apply because fetches are user-initiated; not a training controlBsource
Claude-SearchBot robots.txt tokenrobots-tokenAnthropicvoluntaryglobalTraining (ClaudeBot); already-indexed contentBsource
Claude-User robots.txt tokenrobots-tokenAnthropicvoluntaryglobalTraining (ClaudeBot) and search indexing (Claude-SearchBot) are separate tokensBsource
ClaudeBot robots.txt tokenrobots-tokenAnthropicvoluntaryglobalAnthropic frames the signal as excluding the site's future materials from training; no statement on removing already-collected data; does...Bsource
Cloudflare AI Scrapers and Crawlers blocknetwork-blockCloudflarecontractualglobalOnly sites behind Cloudflare; ML detection is imperfect; user-initiated agents and headless browsers can blur categories; does nothing fo...Bsource
Cloudflare Content Signals Policy (Content-Signal in robots.txt)standardCloudflarevoluntaryglobalCloudflare says signals are preferences, not countermeasures, and some companies may ignore them; ai-input left unset by default; content...Bsource
Cloudflare Pay per crawlnetwork-blockCloudflarecontractualglobalPrivate beta; only sites on Cloudflare; flat domain-wide price, no training vs inference distinction; unregistered crawlers are simply bl...Bsource
EU AI Act Article 53(1)(c)-(d) GPAI copyright policy and training summarylegal-reservationEuropean Unionenforceable-lawEUDoes not itself create a new opt-out; transitional rules for models placed before 2025-08-02 not verified; does not require removal from ...Bsource
EU TDM rights reservation (DSM Directive Article 4(3))legal-reservationEuropean Union (Directive (EU) 2019, 790)enforceable-lawEUResearch organisations under Article 3 (no opt-out); works mined before the reservation was made; training performed outside the EU; uncl...Bsource
GDPR Article 21 right to object (general)regulatory-requestEuropean Union, mirrored in UK GDPRenforceable-lawEUOnly personal data, not copyright in works; controller may continue if it shows compelling legitimate grounds; unlearning from trained mo...Bsource
Google search snippet controls (nosnippet, data-nosnippet, max-snippet, noindex) for AI Overviews and AI Moderobots-tokenGooglevoluntaryglobalCannot opt out of AI features while keeping full Search snippets; recrawl can take days to months; does not govern Gemini training (Googl...Bsource
Google-CloudVertexBot robots.txt tokenrobots-tokenGooglevoluntaryglobalNot a general training opt-out; no effect on Search or other productsBsource
Google-Extended robots.txt tokenrobots-tokenGooglevoluntaryglobalDoes not crawl separately (reuses Googlebot); does NOT affect Google Search, AI Overviews or AI Mode, which are controlled only via Googl...Bsource
GPTBot robots.txt tokenrobots-tokenOpenAIvoluntaryglobalData already collected before the disallow; content reaching OpenAI via third-party datasets or licensed sources; does not affect ChatGPT...Bsource
IETF AI Preferences (aipref) vocabulary and attachmentstandardIETF aipref working groupproposedglobalCharter explicitly excludes technical enforcement; not yet an RFC; no retroactive effectBsource
Known Agents (formerly Dark Visitors) robots.txt agent directorynetwork-blockKnown AgentsvoluntaryglobalCovers crawler identity and robots.txt only; no legal force, platform toggles, regulatory routes, or data-already-collected gapBsource
LinkedIn Data for Generative AI Improvement settingplatform-toggleLinkedIncontractualglobalDoes not affect training that already took place; feedback data and non-generative models need the separate Data Processing Objection formBsource
Meta AI training objection (GDPR Article 21 route)regulatory-requestMeta Platformsenforceable-lawEUPrivate messages and under-18 data already excluded; announcement does not say objection removes data already used; content about you pos...Bsource
Meta-ExternalAgent robots.txt tokenrobots-tokenMetavoluntaryglobalMeta's page does not explicitly state robots.txt compliance for this agent; Meta-ExternalFetcher (user-initiated) may bypass robots.txt; ...Bsource
Meta-ExternalFetcher tokenrobots-tokenMetavoluntaryglobalMeta states it may bypass robots.txtBsource
noai / noimageai directives (meta robots and X-Robots-Tag)standardDeviantArtvoluntaryglobalNo crawler operator documentation read commits to honouring it; DeviantArt acknowledges it cannot technically prevent scraping; content a...Bsource
noarchive meta tag as AI-training signal (Amazon)robots-tokenAmazonvoluntaryglobalOnly Amazon agents documented; page-level onlyBsource
NOCACHE and NOARCHIVE meta tags for Bing Chat / Microsoft generative AIrobots-tokenMicrosoftvoluntaryglobalDefault (no tag) permits training use; announcement dates from 2023 and product names have since changed (Copilot); models already trainedBsource
OAI-SearchBot robots.txt tokenrobots-tokenOpenAIvoluntaryglobalTraining (governed by GPTBot); user-initiated fetches (ChatGPT-User)Bsource
Perplexity-User tokenrobots-tokenPerplexityvoluntaryglobalOperator states it generally ignores robots.txt, so the token is identification only, not an opt-outBsource
PerplexityBot robots.txt tokenrobots-tokenPerplexityvoluntaryglobalPerplexity-User fetches, which Perplexity says generally ignore robots.txt; undeclared crawling observed by CloudflareBsource
RSL (Really Simple Licensing)standardRSL Collective, RSL Internet CollectivevoluntaryglobalNo AI developer commitment to honour it listed on the site; licensing enforcement depends on CDN partners or contractsBsource
Spawning ai.txt and Do Not Train registrystandardSpawningvoluntaryglobalCould not verify current status: spawning.ai showed an under-maintenance page on 2026-09-16; honouring parties not verifiedBsource
Squarespace Block known artificial intelligence crawlersplatform-toggleSquarespacevoluntaryglobalOnly cooperative crawlers; no retroactive removal; no page-level controlBsource
TDM Reservation Protocol (TDMRep)standardW3C TDM Reservation Protocol Community GroupvoluntaryEUIs a W3C Community Group report, not a W3C Standard; no crawler operator documentation read commits to reading it; content mined before p...Bsource
UK TDM opt-out exception (proposed, abandoned)legal-reservationUK GovernmentproposedUKNot law: March 2026 report says a broad exception with opt-out is no longer the preferred way forward; current law (CDPA s29A) permits TD...Bsource
WordPress.com Prevent third-party sharingplatform-toggleAutomatticcontractualglobalRobots.txt part depends on AI platforms honouring it; per-site setting; no statement on data already sharedBsource
YouTube third-party training settingplatform-toggleYouTubecontractualglobalDefault off means no third-party permission, but it does not stop scraping by parties outside the programme; page does not cover Google's...Bsource
Cite as: Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Published by Blomega (Wikidata Q141048865). Unknown values are shown as not established, never guessed.