Blomega Data Refinery / registry
AI benchmarks: data rights, annotator pay and contamination claims
38 records, built 2026-09-18. Machine-readable: JSON and JSON-LD.
| Name | Maintainer | Year | Data licence | Commercial use | Annotator pay | Independent contamination claims | Grade | Source |
|---|---|---|---|---|---|---|---|---|
| AgentBench | Tsinghua University KEG, THUDM | 2023 | Apache-2.0 | not established | not established | 0 | B | source |
| AIME 2025 | Mathematical Association of America, evaluation sets published by unaffiliated Hugging Face uploaders | 2025 | disputed: apache-2.0, mit and no tag on three mirrors of the same 30 problems | not established | not established | 1 | B | source |
| ARC-AGI-2 | ARC Prize Foundation | 2025 | Apache-2.0 | yes | not established | 0 | B | source |
| BFCL (Berkeley Function-Calling Leaderboard) | Gorilla team, UC Berkeley Sky Computing Lab | 2024 | Apache-2.0 (harness repo) / apache-2.0 (card tag, no LICENSE file) | not established | not established | 0 | B | source |
| BIG-bench | 2022 | Apache-2.0 | not established | not established | 0 | B | source | |
| BrowseComp | OpenAI | 2025 | not established | not established | not established | 0 | B | source |
| FrontierMath | Epoch AI | 2024 | proprietary, not released | no | not established | 0 | B | source |
| GPQA | David Rein et al. | 2023 | CC-BY-4.0 (HF card) / MIT (GitHub repo) | yes | $10 base per question plus bonuses ($20 per expert validator who answers correctly, $15 per non-expert who answers incorrectly, $30 extra... | 0 | B | source |
| GPQA Diamond | David Rein et al. | 2023 | CC-BY-4.0 | yes | $10 base per question plus bonuses ($20 per expert validator answering correctly, $15 per non-expert answering incorrectly, $30 extra); e... | 0 | B | source |
| GSM1k | Scale AI | 2024 | not established | not established | not established | 0 | B | source |
| GSM8K | OpenAI | 2021 | MIT | yes | not established | 1 | B | source |
| HellaSwag | Rowan Zellers et al. | 2019 | not established | not established | not established | 1 | B | source |
| HMMT February 2025 (MathArena) | SRI Lab, ETH Zurich, problems by the Harvard-MIT Mathematics Tournament | 2025 | CC-BY-NC-SA-4.0 (card tag) / MIT (harness repo) | no | not established | 0 | B | source |
| HumanEval | OpenAI | 2021 | MIT | yes | not established | 1 | B | source |
| Humanity's Last Exam (HLE) | Center for AI Safety and Scale AI | 2025 | MIT | yes | $500,000 prize pool: $5,000 for each of the top 50 questions, $500 for each of the next 500; plus paper co-authorship for accepted questions | 0 | B | source |
| LIBERO | University of Texas at Austin | 2023 | CC-BY-4.0 | yes | not established | 0 | B | source |
| LIBERO-plus | Senyu Fei, Xipeng Qiu et al. | 2025 | not established | not established | not established | 0 | B | source |
| LiveBench | LiveBench team | 2024 | not established | not established | not established | 0 | B | source |
| LiveCodeBench | Naman Jain et al. | 2024 | MIT (harness repo) / 'cc' unspecified (dataset card) | not established | not established | 1 | B | source |
| MATH | Dan Hendrycks et al. | 2021 | MIT | not established | not established | 1 | B | source |
| MBPP (Mostly Basic Python Problems) | Google Research | 2021 | CC-BY-4.0 | not established | not established | 1 | B | source |
| MMLU | Dan Hendrycks et al. | 2020 | MIT | not established | not established | 2 | B | source |
| MMLU-Pro | TIGER-Lab, University of Waterloo | 2024 | mit (card tag only) | not established | not established | 0 | B | source |
| MMMU | Xiang Yue et al. | 2023 | Apache-2.0 | not established | not established | 0 | B | source |
| MTEB (Massive Text Embedding Benchmark) | MTEB community, originally Niklas Muennighoff, Nouamane Tazi, Loic Magne and Nils Reimers | 2022 | Apache-2.0 (framework only) | not established | not established | 0 | B | source |
| OSWorld | XLANG Lab, University of Hong Kong | 2024 | Apache-2.0 | not established | not established | 0 | B | source |
| RewardBench | Allen Institute for AI | 2024 | ODC-BY | not established | not established | 0 | B | source |
| RoboArena | Pranav Atreya, Karl Pertsch et al. | 2025 | MIT | not established | not established | 0 | B | source |
| RoboCasa | Soroush Nasiriany, Ajay Mandlekar, Yuke Zhu et al. | 2024 | MIT | not established | not established | 0 | B | source |
| SimpleQA | OpenAI | 2024 | MIT | yes | not established | 0 | B | source |
| SimplerEnv | Xuanlin Li, Kyle Hsu et al. | 2024 | MIT | not established | not established | 0 | B | source |
| SWE-bench | Carlos E. Jimenez, John Yang et al. | 2023 | MIT (harness repo) / none declared (dataset repo) | not established | not established | 3 | B | source |
| SWE-bench Verified | SWE-bench team (Princeton) with OpenAI Preparedness | 2024 | not established | not established | not established | 1 | B | source |
| tau-bench | Sierra AI | 2024 | MIT | not established | not established | 0 | B | source |
| Terminal-Bench | Laude Institute with Stanford University | 2025 | Apache-2.0 | not established | not established | 0 | B | source |
| TruthfulQA | Stephanie Lin, Jacob Hilton, Owain Evans | 2021 | Apache-2.0 | yes | not established | 1 | B | source |
| WebArena | Carnegie Mellon University | 2023 | Apache-2.0 | not established | not established | 0 | B | source |
| WinoGrande | Allen Institute for AI | 2019 | CC-BY (version unstated) | yes | $0.40 per twin-sentence pair written; $0.03 per sentence validation (paper footnotes 4 and 5) | 2 | B | source |
Cite as: Blomega Data Refinery, https://data.blomega.com, CC BY 4.0. Published by
Blomega (Wikidata Q141048865). Unknown values are shown as
not established, never guessed.