Blomega Data Refinery / registry

AI benchmarks: data rights, annotator pay and contamination claims

38 records, built 2026-09-18. Machine-readable: JSON and JSON-LD.

NameMaintainerYearData licenceCommercial useAnnotator payIndependent contamination claimsGradeSource
AgentBenchTsinghua University KEG, THUDM2023Apache-2.0not establishednot established0Bsource
AIME 2025Mathematical Association of America, evaluation sets published by unaffiliated Hugging Face uploaders2025disputed: apache-2.0, mit and no tag on three mirrors of the same 30 problemsnot establishednot established1Bsource
ARC-AGI-2ARC Prize Foundation2025Apache-2.0yesnot established0Bsource
BFCL (Berkeley Function-Calling Leaderboard)Gorilla team, UC Berkeley Sky Computing Lab2024Apache-2.0 (harness repo) / apache-2.0 (card tag, no LICENSE file)not establishednot established0Bsource
BIG-benchGoogle2022Apache-2.0not establishednot established0Bsource
BrowseCompOpenAI2025not establishednot establishednot established0Bsource
FrontierMathEpoch AI2024proprietary, not releasednonot established0Bsource
GPQADavid Rein et al.2023CC-BY-4.0 (HF card) / MIT (GitHub repo)yes$10 base per question plus bonuses ($20 per expert validator who answers correctly, $15 per non-expert who answers incorrectly, $30 extra...0Bsource
GPQA DiamondDavid Rein et al.2023CC-BY-4.0yes$10 base per question plus bonuses ($20 per expert validator answering correctly, $15 per non-expert answering incorrectly, $30 extra); e...0Bsource
GSM1kScale AI2024not establishednot establishednot established0Bsource
GSM8KOpenAI2021MITyesnot established1Bsource
HellaSwagRowan Zellers et al.2019not establishednot establishednot established1Bsource
HMMT February 2025 (MathArena)SRI Lab, ETH Zurich, problems by the Harvard-MIT Mathematics Tournament2025CC-BY-NC-SA-4.0 (card tag) / MIT (harness repo)nonot established0Bsource
HumanEvalOpenAI2021MITyesnot established1Bsource
Humanity's Last Exam (HLE)Center for AI Safety and Scale AI2025MITyes$500,000 prize pool: $5,000 for each of the top 50 questions, $500 for each of the next 500; plus paper co-authorship for accepted questions0Bsource
LIBEROUniversity of Texas at Austin2023CC-BY-4.0yesnot established0Bsource
LIBERO-plusSenyu Fei, Xipeng Qiu et al.2025not establishednot establishednot established0Bsource
LiveBenchLiveBench team2024not establishednot establishednot established0Bsource
LiveCodeBenchNaman Jain et al.2024MIT (harness repo) / 'cc' unspecified (dataset card)not establishednot established1Bsource
MATHDan Hendrycks et al.2021MITnot establishednot established1Bsource
MBPP (Mostly Basic Python Problems)Google Research2021CC-BY-4.0not establishednot established1Bsource
MMLUDan Hendrycks et al.2020MITnot establishednot established2Bsource
MMLU-ProTIGER-Lab, University of Waterloo2024mit (card tag only)not establishednot established0Bsource
MMMUXiang Yue et al.2023Apache-2.0not establishednot established0Bsource
MTEB (Massive Text Embedding Benchmark)MTEB community, originally Niklas Muennighoff, Nouamane Tazi, Loic Magne and Nils Reimers2022Apache-2.0 (framework only)not establishednot established0Bsource
OSWorldXLANG Lab, University of Hong Kong2024Apache-2.0not establishednot established0Bsource
RewardBenchAllen Institute for AI2024ODC-BYnot establishednot established0Bsource
RoboArenaPranav Atreya, Karl Pertsch et al.2025MITnot establishednot established0Bsource
RoboCasaSoroush Nasiriany, Ajay Mandlekar, Yuke Zhu et al.2024MITnot establishednot established0Bsource
SimpleQAOpenAI2024MITyesnot established0Bsource
SimplerEnvXuanlin Li, Kyle Hsu et al.2024MITnot establishednot established0Bsource
SWE-benchCarlos E. Jimenez, John Yang et al.2023MIT (harness repo) / none declared (dataset repo)not establishednot established3Bsource
SWE-bench VerifiedSWE-bench team (Princeton) with OpenAI Preparedness2024not establishednot establishednot established1Bsource
tau-benchSierra AI2024MITnot establishednot established0Bsource
Terminal-BenchLaude Institute with Stanford University2025Apache-2.0not establishednot established0Bsource
TruthfulQAStephanie Lin, Jacob Hilton, Owain Evans2021Apache-2.0yesnot established1Bsource
WebArenaCarnegie Mellon University2023Apache-2.0not establishednot established0Bsource
WinoGrandeAllen Institute for AI2019CC-BY (version unstated)yes$0.40 per twin-sentence pair written; $0.03 per sentence validation (paper footnotes 4 and 5)2Bsource
Cite as: Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Published by Blomega (Wikidata Q141048865). Unknown values are shown as not established, never guessed.