Dossier 10
Measurement
v1.6-draft · unchecked
- E1 confirmed
- E2 single-source
- E3 claimed
- E4 disputed
- E0 unknown
1. As of
- Date
- Version
- v1.6-draft
- Author / model
- Atlas generator
- Reviewer
- unchecked (Josef)
2. In one sentence
Progress on leaderboards is progress on an instrument: HELM shows that accuracy hides other desiderata; Arena measures preference, not truth; SWE-bench measures tests, not software engineering; LiveCodeBench shows that time windows are the hardest contamination control — and only for contest problems; HumanEval (Chen 2021, E1) measures functional correctness on 164 handwritten Python functions (Codex-12B pass@1 ≈28.8%), not multi-file SWE; GPQA measures hard Google-proof Q&A; GSM8K and MATH are math instruments; BBH shows answer-only underestimates CoT-capable multi-step tasks — paper snapshots, not 2026 live boards.
Established now · E1 / E2
3. What works today
Benchmarks are instruments, not thermometers
-
E1
HELM taxonomizes scenarios (task × domain × language) and measures 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, efficiency) on 16 core scenarios, 87.5% of 112 pairs (98/112). 30 models, 42 scenarios, 4,939 runs. Before: models evaluated on 17.9% of core scenarios on average; after 96.0%. E1 for the construct and 2022 campaign; E0 for 2026 frontier scores.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper. -
E1
Landing 2026: Classic, Lite, Capabilities, Instruct, Audio, VHELM, Safety, Finance, MedHELM, Long Context and others exist. E1 for the instrument ecosystem; result tables not opened.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.HELM landing. Holistic Evaluation of Language Models (HELM) Leaderboards. https://crfm.stanford.edu/helm/. As of page as of 2026-08-28. Checked 2026-08-28. Type: Official site. -
E1
What accuracy hides (HELM findings 2022): instruction-tuning leads accuracy/robustness/fairness despite 10× smaller scale. Accuracy–calibration is scenario-dependent. NarrativeQA TNLG v2 72.6% → 38.9% under robustness perturbations. Prompting: wild swings 30% → 80%. Multiple-choice adaptation: OPT 175B HellaSwag 79.1% (separate 0-shot) vs. 30.2% (joint 5-shot). Scale predicts accuracy within families; across families for no scenario. E1 for the 2022 snapshot.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.
Four instruments: construct and limit
-
E2
Arena: anonymous side-by-side battles, Bradley-Terry MLE (not online Elo). Paper Jan 2024: ~240k votes, ~90k users, >50 models. Crowd–expert agreement 72.8–83.1%. User base primarily LLM hobbyists; helpfulness, not safety. Live ranks 2026: E0 (board not opened).
Chiang et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. As of 2024-03-07 (ICML 2024). Checked 2026-08-28. Type: Paper. -
E2
Style control Aug 2024: length coefficient 0.249 — dominant. After control: GPT-4o-mini and Grok-2-mini fall; Claude 3.5 Sonnet/Opus, Llama-3.1-405B rise. Observational, not causal. Length is part of human preference, not only noise.
Li, Angelopoulos, Chiang. Does style matter? Disentangling style and substance in Chatbot Arena. https://www.lmsys.org/blog/2024-08-28-style-control/. As of 2024-08-29. Checked 2026-08-28. Type: Official blog. -
E1
SWE-bench: 2,294 issues, 12 Python repos. % resolved = all FAIL_TO_PASS and PASS_TO_PASS green. Claude 2 + BM25: 1.96% resolved; oracle retrieval 4.80%. Apply rates ≫ resolve. Longer context lowers resolve (1.96 → 1.22% at 50k). SWE-Verified / Pro / + E0 in this log (OpenAI timeout).
Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper. -
E2
LiveCodeBench arXiv HTML: 511 problems May 2023–May 2024. Time window: DeepSeek-Instruct-33B drops after Aug 2023; GPT-4o after Nov 2023; Codestral 36.5% → 28.3%. HumanEval overfitting: DS-Ins-1.3B 59.8% HumanEval+ vs. 26.3% LCB-Easy. GPT-4-Turbo-2024-04-09 code-gen total 41.1 Pass@1; hard 5.1. Live board 2026 E0.
Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper. -
E1
Fact sheets — what is measured, what is not.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.Chiang et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. As of 2024-03-07 (ICML 2024). Checked 2026-08-28. Type: Paper.Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper.Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper.Four instruments. Construct and opened limits only. No 2026 ranking. Instrument Measures Does not Noise HELM Scenarios × 7 metrics, 5-shot Agents, non-EN core, CoT 2022 Prompt/MC, auto-summarization Arena Pairwise human preference, live Safety, facts, execution Length/markdown, no GT SWE-bench Patch per repo tests Design, review, multi-language Retrieval, apply, 12 repos LiveCodeBench Contest code, time-filtered Repo SE, other languages n after cutoff, LeetCode leak
Truth vs preference: TruthfulQA
-
E1
TruthfulQA = 817 questions, 38 categories, zero-shot, adversarial against imitative falsehoods. Human (internet, ~2 min/question): 94% true, 87% true+informative. Best model in the paper (GPT-3-175B, helpful prompt): 58% true, 21% true+informative; 42% false+informative (human 6%). Inverse scaling: larger models in the same family often less truthful (GPT-Neo/J 6B 17% less truthful than a 60× smaller one). UnifiedQA more truthful but less informative. Multiple-choice: no model significantly above chance; larger often worse.
Lin, Hilton, Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. https://arxiv.org/html/2109.07958. As of 2021-09-08. Checked 2026-08-29. Type: Paper (NeurIPS). -
E2
Inverse scaling is explained in the paper as imitative falsehoods (better learning of the training distribution), with control questions (edits) that improve with size, and paraphrases that hold the trend. That is 2021 GPT-3/J/2/T5, not the 2026 frontier. Appendix B.3 (newer Anthropic/InstructGPT/WebGPT/Gopher models) shows progress and a return to positive scaling at the largest, but still a gap to humans — E2 for these appendix figures (external eval, opened in the same HTML).
Lin, Hilton, Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. https://arxiv.org/html/2109.07958. As of 2021-09-08. Checked 2026-08-29. Type: Paper (NeurIPS).
MMLU: saturation is ceiling plus noise
-
E1
MMLU: 57 tasks, 15,908 questions. GPT-3 175B few-shot 43.9%. MTurk 34.5%; expert estimate ~89.8%. Entropy–accuracy negatively correlated (r=−0.43 zero-shot) — no signal for verbatim Q&A memorization in the GPT-3 era.
Hendrycks et al. Measuring Massive Multitask Language Understanding. https://arxiv.org/html/2009.03300. As of 2020-09 / ICLR 2021. Checked 2026-08-28. Type: Paper. -
E2
MMLU-Pro: frontier clusters on MMLU 86–87% (GPT-4 86.4% as a citation; GPT-4 report not opened here). MMLU-Pro: 12,032 questions, mean 9.47 options. GPT-4o 72.6% CoT vs. 53.5% direct. Gap GPT-4o vs. GPT-4-Turbo ~1 pp MMLU → ~9 pp MMLU-Pro. Prompt sensitivity peak 10.98% → 3.74%. Saturation ≠ expert parity.
Wang et al. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. https://arxiv.org/html/2406.01574. As of 2024-06 / NeurIPS 2024. Checked 2026-08-28. Type: Paper.
Contamination: mem ≠ expl; little verbatim on four open models
-
E1
Magar: mem (MLM seen−unseen) ≠ expl (task seen−unseen). SST-5, 200 duplicates: mem ~60%, expl ~40%. mem 15% can yield expl <1%. Early contamination → high expl; late + LR-decay → expl ~0 with still-high mem. BERT scale, not frontier LLM.
Magar & Schwartz. Data Contamination: From Memorization to Exploitation. https://arxiv.org/html/2203.08242. As of 2022-03 / ACL 2022. Checked 2026-08-28. Type: Paper. -
E1
Oren: exchangeability test, FPR guarantee. Audit Llama2-7B, Mistral-7B, Pythia-1.4B, GPT-2 XL: no significant verbatim contamination (exception Mistral × ARC-Easy). Limit: order memorization only, not paraphrases. E1 method, E2 audit.
Oren et al. Proving Test Set Contamination in Black-Box Language Models. https://arxiv.org/html/2310.17623. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper. -
E2
Sainz: guideline / raw-text / annotation contamination; after deployment (API logs). ChatGPT/WizardCoder regenerate CoNLL-2003 train start (appendix). GSM8K body now opened as source 18 (E1 paper); GPT-4 figures from the Technical Report in Sainz remain E3 second-order (report not opened here).
Sainz et al. NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. https://arxiv.org/html/2310.18018. As of 2023-10 / Findings EMNLP 2023. Checked 2026-08-28. Type: Position paper.
Metric artefacts and unstable ranks
-
E1
Schaeffer: nonlinear metrics (accuracy, exact match, multiple choice grade) produce sharpness from smooth per-token error. >92% of hand-annotated BIG-Bench emergences under MC grade or exact string match. “nothing … cannot display emergent abilities” — only the published claims are suspect. Not a refutation of every possible emergence.
Schaeffer et al. Are Emergent Abilities of Large Language Models a Mirage?. https://arxiv.org/html/2304.15004. As of 2023-04 / NeurIPS 2023. Checked 2026-08-28. Type: Paper. -
E1
Wei et al. (arXiv:2206.07682, TMLR 2022, body opened): emergence definition — an ability is emergent if absent in smaller models but present in larger ones; not predictable by extrapolating a scaling curve (near-random until a threshold, then a phase transition).
Wei et al. Emergent. Emergent Abilities of Large Language Models. https://arxiv.org/pdf/2206.07682. As of 2022-06 / TMLR 2022. Checked 2026-09-06. Type: Paper (arXiv; TMLR). -
E1
Wei Table 1 / Fig. 2: few-shot examples — e.g. 3-digit arithmetic ~2×10²² FLOPs / 13B GPT-3; MMLU ~3×10²³ / 175B GPT-3; WiC only at ~2.5×10²⁴ / 540B PaLM. BIG-Bench: >200 tasks; dozens still near-random even at largest GPT-3/PaLM (paper).
Wei et al. Emergent. Emergent Abilities of Large Language Models. https://arxiv.org/pdf/2206.07682. As of 2022-06 / TMLR 2022. Checked 2026-09-06. Type: Paper (arXiv; TMLR). -
E1
Wei limits the claim: no fixed universal threshold; scale not the only factor; not a proof that all abilities emerge. Survey/definition — contrast with Schaeffer stays open; no reconciliation here.
Wei et al. Emergent. Emergent Abilities of Large Language Models. https://arxiv.org/pdf/2206.07682. As of 2022-06 / TMLR 2022. Checked 2026-09-06. Type: Paper (arXiv; TMLR). -
E2
Alzahrani: mini-perturbations on MMLU shift ranks by up to 8 places (11 models). Llama-2-7b zero-shot: correct-at-A 66.36% vs. correct-at-D 23.37% (Δ 43 pp). In-context cheating: 5-shot with correct answer 92–99%; with wrong → collapse.
Alzahrani et al. When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. https://arxiv.org/html/2402.01781. As of 2024-02. Checked 2026-08-28. Type: Paper. -
E2
Bowman & Dahl 2021: GLUE/SuperGLUE saturates while failing CheckList/McCoy. 80→81% needs thousands of items. Saturation as a design fault, not mastery.
Bowman & Dahl. What Will it Take to Fix Benchmarking in Natural Language Understanding?. https://arxiv.org/html/2104.02145. As of 2021-04 / NAACL 2021. Checked 2026-08-28. Type: Position paper.
Graduate-level Google-proof Q&A: GPQA
-
E1
GPQA (Rein et al., arXiv:2311.12022, body opened): 448 expert-written MCQs (main set) in biology/physics/chemistry; Extended 546; Diamond 198.
Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM). -
E1
Experts (PhD domain) ~65% accuracy on Extended (74% when discounting clear mistakes / post-hoc agreement analysis).
Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM). -
E1
Skilled non-experts with web access ~34% (avg ≥30 min); questions designed to be Google-proof.
Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM). -
E1
Strongest GPT-4 few-shot CoT baseline ~39% (chance 25%).
Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM). -
E1
Diamond: both experts agree and ≤1/3 non-experts correct (highest-quality subset).
Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM).
Claimed · E3
4. What is claimed, not shown
Snapshots 2024, not frontier 2026
-
E2
MMLU-Pro CoT May 2024: GPT-4o 72.6; Gemini-1.5-Pro 69.0; Claude-3-Opus 68.5; GPT-4-Turbo 63.7; Llama-3-70B-Instruct 56.2. One eval harness.
Wang et al. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. https://arxiv.org/html/2406.01574. As of 2024-06 / NeurIPS 2024. Checked 2026-08-28. Type: Paper. -
E2
Style-control rank shifts Aug 2024 (chatgpt-4o-latest stays #1; grok-2-mini 6→18) are one blog snapshot, not 2026.
Li, Angelopoulos, Chiang. Does style matter? Disentangling style and substance in Chatbot Arena. https://www.lmsys.org/blog/2024-08-28-style-control/. As of 2024-08-29. Checked 2026-08-28. Type: Official blog. -
E2
HELM 2022 winners: text-davinci-002, Anthropic-LM 52B, TNLG v2 530B. 2022 models ≠ 2026 frontier.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.
Math instruments: GSM8K and MATH
-
E1
GSM8K (Cobbe et al., arXiv:2110.14168, body opened): 8.5K human-written grade-school math word problems; 7.5K train / 1K test; 2–8 steps; elementary arithmetic (+−×÷); <2% estimated breaking errors. Calculator annotations at train/test.
Cobbe et al. GSM8K. Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. As of 2021-11. Checked 2026-09-08. Type: Paper (arXiv). -
E1
6B finetuned with intermediate steps: 20.6% test solve; without intermediate steps: 5.2%. Verification (ranking many completions) improves substantially over finetuning and scales better with more data than the finetuning baseline. Naive log-linear extrapolation in the paper: ~10¹⁶ parameters for 80% solve on full GSM8K train.
Cobbe et al. GSM8K. Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. As of 2021-11. Checked 2026-09-08. Type: Paper (arXiv). -
E1
MATH (Hendrycks et al., arXiv:2103.03874, body opened): 12,500 competition math problems; 7,500 train / 5,000 test; 7 subjects; difficulty 1–5 (AoPS). Large LMs in the paper: 3.0%–6.9% accuracy; easiest level up to ~15%. Humans: CS PhD ~40%; three-time IMO gold ~90%.
Hendrycks et al. MATH. Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. As of 2021-03. Checked 2026-09-08. Type: Paper (arXiv). -
E1
GPT-2 1.5B + AMPS pretrain: 6.9% overall (Table 2); 0.1B: 5.4%. AMPS: >100k Khan + >5M Mathematica, 23 GB. Training on step-by-step raises relative accuracy by 10% vs Q&A only; generating solutions at test time decreased accuracy vs immediate answer. Scaling extrapolation: ~10³⁵ params for 40% (impractical).
Hendrycks et al. MATH. Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. As of 2021-03. Checked 2026-09-08. Type: Paper (arXiv). -
E1
Two instruments, not one: GSM8K measures grade-school word problems; MATH measures competition math. 2021 paper snapshots — no claim about 2026 frontier scores.
Cobbe et al. GSM8K. Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. As of 2021-11. Checked 2026-09-08. Type: Paper (arXiv).Hendrycks et al. MATH. Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. As of 2021-03. Checked 2026-09-08. Type: Paper (arXiv).
BBH: answer-only underestimates multi-step
-
E1
BBH (Suzgun et al., arXiv:2210.09261, body opened): 23 BIG-Bench tasks where prior LM evals did not beat the average human-rater. Avg human-rater BBH all 67.7; max human-rater 94.4; best prior BIG-Bench 50.9 (0/23 above avg).
Suzgun et al. BBH. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. https://arxiv.org/pdf/2210.09261. As of 2022-10 / Findings ACL 2023. Checked 2026-09-09. Type: Paper (arXiv; ACL Findings). -
E1
Table 2: Codex (code-davinci-002) CoT 73.9 (+16.7 vs answer-only 56.6) — 17/23 above avg human (answer-only 5/23). PaLM 540B CoT 65.2 (10/23); InstructGPT text-davinci-002 CoT 68.4 (15/23).
Suzgun et al. BBH. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. https://arxiv.org/pdf/2210.09261. As of 2022-10 / Findings ACL 2023. Checked 2026-09-09. Type: Paper (arXiv; ACL Findings). -
E1
Authors: few-shot without CoT systematically underestimates multi-step BBH; CoT can unlock emergent task performance on otherwise flat scaling curves. Codex CoT > avg human by >6%, but >20% behind max human-rater.
Suzgun et al. BBH. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. https://arxiv.org/pdf/2210.09261. As of 2022-10 / Findings ACL 2023. Checked 2026-09-09. Type: Paper (arXiv; ACL Findings).
HumanEval: functional correctness as an instrument
-
E1
HumanEval (Chen et al., arXiv:2107.03374; body opened 2026-09-12): 164 handwritten Python problems from docstrings; mean ~7.7 unit tests; functional correctness via unit tests (not BLEU). pass@k as an unbiased estimator with n≥k (paper: n=200, k≤100).
Chen et al. HumanEval. Evaluating Large Language Models Trained on Code. https://arxiv.org/pdf/2107.03374. As of arXiv 2021-07. Checked 2026-09-12. Type: Paper (arXiv). -
E1
Codex-12B pass@1 ≈28.8% (abstract/Fig.; Table often 28.81); GPT-3 ≈0%; GPT-J 6B pass@1 ≈11.4–11.62%. With 100 samples: ~70.2% of problems solved (abstract); Codex-S supervised pass@1 37.7%, pass@100 with oracle selection 77.5%. BLEU distributions of correct vs incorrect samples overlap — BLEU ≠ functional correctness.
Chen et al. HumanEval. Evaluating Large Language Models Trained on Code. https://arxiv.org/pdf/2107.03374. As of arXiv 2021-07. Checked 2026-09-12. Type: Paper (arXiv). -
E0
Limit: interview-like single functions, not multi-file/SWE. 2026 live boards and contamination of the 164 tasks not re-measured here (E0).
Chen et al. HumanEval. Evaluating Large Language Models Trained on Code. https://arxiv.org/pdf/2107.03374. As of arXiv 2021-07. Checked 2026-09-12. Type: Paper (arXiv).
Constrained · Limit
5. Bottleneck and limit
What a score is not
-
E1
Leaderboard live 2026 frontier scores on GPQA not opened here (E0 for 2026 live boards). GPQA measures hard Q&A for scalable oversight, not software engineering (contrast SWE-bench, already in the dossier).
Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM). -
E0
2026 live leaderboards for GSM8K and MATH not opened here (E0).
Cobbe et al. GSM8K. Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. As of 2021-11. Checked 2026-09-08. Type: Paper (arXiv).Hendrycks et al. MATH. Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. As of 2021-03. Checked 2026-09-08. Type: Paper (arXiv). -
E0
2026 live leaderboards / replications for BBH not opened here (E0).
Suzgun et al. BBH. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. https://arxiv.org/pdf/2210.09261. As of 2022-10 / Findings ACL 2023. Checked 2026-09-09. Type: Paper (arXiv; ACL Findings). -
E1
HumanEval ≠ SWE-bench / multi-file engineering: single functions from docstrings; pass@k with oracle ≠ deployment pass@1. 2021 paper snapshot ≠ 2026 live boards (E0 for live scores).
Chen et al. HumanEval. Evaluating Large Language Models Trained on Code. https://arxiv.org/pdf/2107.03374. As of arXiv 2021-07. Checked 2026-09-12. Type: Paper (arXiv). -
E1
BBH avg human-rater ≠ max human-rater; CoT score ≠ “reasoning solved”. 2022/23 paper snapshot ≠ 2026 frontier.
Suzgun et al. BBH. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. https://arxiv.org/pdf/2210.09261. As of 2022-10 / Findings ACL 2023. Checked 2026-09-09. Type: Paper (arXiv; ACL Findings). -
E1
GSM8K grade-school ≠ competition MATH ≠ GPQA graduate Q&A — three different instruments; one score does not replace the others.
Cobbe et al. GSM8K. Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. As of 2021-11. Checked 2026-09-08. Type: Paper (arXiv).Hendrycks et al. MATH. Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. As of 2021-03. Checked 2026-09-08. Type: Paper (arXiv).Rein et al. GPQA. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. As of 2023-11-20. Checked 2026-09-05. Type: Paper (arXiv; COLM). -
E1
Paper snapshots (Cobbe 2021 / Hendrycks 2021) ≠ 2026 frontier — E1 only for opened paper measurements.
Cobbe et al. GSM8K. Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. As of 2021-11. Checked 2026-09-08. Type: Paper (arXiv).Hendrycks et al. MATH. Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. As of 2021-03. Checked 2026-09-08. Type: Paper (arXiv). -
E2
Saturation ≠ mastery: MMLU plateau 86–87% with MMLU-Pro 72.6% (27 pp of ceiling); HELM summarization metrics do not discriminate; Bowman: score near ceiling with CheckList fail.
Wang et al. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. https://arxiv.org/html/2406.01574. As of 2024-06 / NeurIPS 2024. Checked 2026-08-28. Type: Paper.Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.Bowman & Dahl. What Will it Take to Fix Benchmarking in Natural Language Understanding?. https://arxiv.org/html/2104.02145. As of 2021-04 / NAACL 2021. Checked 2026-08-28. Type: Position paper. -
E1
Contamination does not explain all gains: Magar expl only under duplication/timing; Oren little verbatim on 4 open models; Hendrycks no entropy coupling. LCB time-window drops can coexist — do not net them against each other.
Magar & Schwartz. Data Contamination: From Memorization to Exploitation. https://arxiv.org/html/2203.08242. As of 2022-03 / ACL 2022. Checked 2026-08-28. Type: Paper.Oren et al. Proving Test Set Contamination in Black-Box Language Models. https://arxiv.org/html/2310.17623. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper.Hendrycks et al. Measuring Massive Multitask Language Understanding. https://arxiv.org/html/2009.03300. As of 2020-09 / ICLR 2021. Checked 2026-08-28. Type: Paper.Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper. -
E2
Arena ≠ capability: pairwise pref, diverse prompts, BT CIs. Not: safety, facts, production, ground truth. Style control is an interpretation, not “true” strength.
Chiang et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. As of 2024-03-07 (ICML 2024). Checked 2026-08-28. Type: Paper.Li, Angelopoulos, Chiang. Does style matter? Disentangling style and substance in Chatbot Arena. https://www.lmsys.org/blog/2024-08-28-style-control/. As of 2024-08-29. Checked 2026-08-28. Type: Official blog. -
E1
Arena measures preference (helpfulness); TruthfulQA measures avoidance of imitative falsehood. These are different instruments. An Arena rank is not a truth rank.
Chiang et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. As of 2024-03-07 (ICML 2024). Checked 2026-08-28. Type: Paper.Lin, Hilton, Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. https://arxiv.org/html/2109.07958. As of 2021-09-08. Checked 2026-08-29. Type: Paper (NeurIPS). -
E1
SWE % resolved ≠ SE ability: execution-based, real issues. Not: retrieval-independent, multi-language, design. 2023 SOTA 2%.
Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper. -
E1
Emergence not refuted: many claims are metric-induced. Schaeffer does not rule out real emergence. HELM finding 25 (scale threshold ≥50B on accuracy) remains compatible.
Schaeffer et al. Are Emergent Abilities of Large Language Models a Mirage?. https://arxiv.org/html/2304.15004. As of 2023-04 / NeurIPS 2023. Checked 2026-08-28. Type: Paper.Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper. -
E0
SWE-Verified dead as a frontier measure: OpenAI page timeout. 59.4% flawed tests = E0. Model names from aggregator snippets (GPT-5.2, Claude Opus 4.5, Gemini 3 Flash) not in the log.
6. Actors and incentives
Who builds the instrument
-
E1
Stanford CRFM runs HELM as a living benchmark; the 2022 paper is the campaign, the 2026 landing is the ecosystem.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.HELM landing. Holistic Evaluation of Language Models (HELM) Leaderboards. https://crfm.stanford.edu/helm/. As of page as of 2026-08-28. Checked 2026-08-28. Type: Official site. -
E2
LMSYS / LMArena: preference board plus style control. Live 2026 not opened.
Chiang et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. As of 2024-03-07 (ICML 2024). Checked 2026-08-28. Type: Paper.Li, Angelopoulos, Chiang. Does style matter? Disentangling style and substance in Chatbot Arena. https://www.lmsys.org/blog/2024-08-28-style-control/. As of 2024-08-29. Checked 2026-08-28. Type: Official blog. -
E1
Princeton/NLP group (SWE-bench) and Berkeley/MIT et al. (LiveCodeBench) supply execution evals with their own confounds.
Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper.Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper.
7. State of the dispute
Emergence, contamination, one score
-
E1
Schaeffer vs. Wei: both bodies opened (sources 7 and 17). Wei: unpredictable scale thresholds; Schaeffer: many claims metric-induced. No reconciliation in this entry. HELM scale threshold remains.
Schaeffer et al. Are Emergent Abilities of Large Language Models a Mirage?. https://arxiv.org/html/2304.15004. As of 2023-04 / NeurIPS 2023. Checked 2026-08-28. Type: Paper.Wei et al. Emergent. Emergent Abilities of Large Language Models. https://arxiv.org/pdf/2206.07682. As of 2022-06 / TMLR 2022. Checked 2026-09-06. Type: Paper (arXiv; TMLR).Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper. -
E2
Oren “little evidence for pervasive contamination” on four open models vs. LCB performance drops on post-cutoff contest problems. Both can be true.
Oren et al. Proving Test Set Contamination in Black-Box Language Models. https://arxiv.org/html/2310.17623. As of 2023-10 / ICLR 2024. Checked 2026-08-28. Type: Paper.Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper.
8. Open questions
- OpenAI SWE-Verified withdrawal 2026-02-23 fetch again — timeout.
- SWE-Bench Pro (Deng 2025) and SWE-bench+ (Aleithan) not opened in this run.
- Chen HumanEval 2021 opened 2026-09-12 (source 21, E1); Liu HumanEval+ 2023 still unopened.
- Wei 2022 Emergent Abilities — opened 2026-09-06 (source 17).
- GSM8K (Cobbe) and MATH (Hendrycks) — opened 2026-09-08 (sources 18–19); BBH (Suzgun) opened 2026-09-09 (source 20); 2026 live boards still E0.
- HLE, LiveBench, SimpleQA; GPQA body opened 2026-09-05 (source 16); 2026 live boards still E0.
- HELM Classic/Capabilities live tables 2026.
- Arena.ai / LMArena live board 2026.
- LiveCodeBench leaderboard + ICLR 2025 (600+ problems).
- GPT-4 Technical Report (MMLU 86.4% only via Wang).
9. Changes
- v1.6-draft2026-09-12: HumanEval (Chen et al., arXiv:2107.03374, source 21, E1) opened — 164 problems, ~7.7 tests; Codex-12B pass@1 ≈28.8%; GPT-J ≈11.4–11.62%; Codex-S 37.7% / oracle pass@100 77.5%; BLEU ≠ correctness; live boards/contamination E0.
- v1.5-draft2026-09-09: BBH (Suzgun et al., arXiv:2210.09261, source 20, E1) opened — answer-only underestimates multi-step; Codex CoT 73.9 / 17 of 23 above avg human; live boards E0.
- v1.4-draft2026-09-08: GSM8K (Cobbe et al., arXiv:2110.14168, source 18, E1) and MATH (Hendrycks et al., arXiv:2103.03874, source 19, E1) opened — math instruments; live boards 2026 E0; Sainz note on GSM8K body updated.
- v1.3-draft2026-09-06: Wei et al. Emergent Abilities (arXiv:2206.07682, source 17, E1) opened — definition + Table 1 examples; contrast with Schaeffer without reconciliation.
- v1.2-draft2026-09-05: GPQA (Rein et al., arXiv:2311.12022 PDF, source 16) opened — 448/546/198; experts ~65% (74% post-hoc); non-experts ~34%; GPT-4 few-shot CoT ~39%; Diamond filter. Live boards 2026 still E0.
- v1.1-draft2026-08-29: TruthfulQA (arXiv:2109.07958 HTML) opened; inverse scaling 2021 as an instrument, not as 2026 SOTA.
- v1.0-draftFirst version from the 2026-08-28 verification log.
10. Sources
| No. | Source | As of | Checked | Grade |
|---|---|---|---|---|
| 1 | . Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. Type: Paper. | E1 | ||
| 2 | . Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. Type: Paper. | E2 | ||
| 3 | . SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. Type: Paper. | E1 | ||
| 4 | . LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. Type: Paper. | E2 | ||
| 5 | . Measuring Massive Multitask Language Understanding. https://arxiv.org/html/2009.03300. Type: Paper. | E1 | ||
| 6 | . MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. https://arxiv.org/html/2406.01574. Type: Paper. | E2 | ||
| 7 | . Are Emergent Abilities of Large Language Models a Mirage?. https://arxiv.org/html/2304.15004. Type: Paper. | E1 | ||
| 8 | . Data Contamination: From Memorization to Exploitation. https://arxiv.org/html/2203.08242. Type: Paper. | E1 | ||
| 9 | . NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. https://arxiv.org/html/2310.18018. Type: Position paper. | E2 | ||
| 10 | . Proving Test Set Contamination in Black-Box Language Models. https://arxiv.org/html/2310.17623. Type: Paper. | E1 | ||
| 11 | . What Will it Take to Fix Benchmarking in Natural Language Understanding?. https://arxiv.org/html/2104.02145. Type: Position paper. | E2 | ||
| 12 | . When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. https://arxiv.org/html/2402.01781. Type: Paper. | E2 | ||
| 13 | . Does style matter? Disentangling style and substance in Chatbot Arena. https://www.lmsys.org/blog/2024-08-28-style-control/. Type: Official blog. | E2 | ||
| 14 | . Holistic Evaluation of Language Models (HELM) Leaderboards. https://crfm.stanford.edu/helm/. Type: Official site. | E1 | ||
| 15 | . TruthfulQA: Measuring How Models Mimic Human Falsehoods. https://arxiv.org/html/2109.07958. Type: Paper (NeurIPS). | E1 | ||
| 16 | . GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/pdf/2311.12022. Type: Paper (arXiv; COLM). | E1 | ||
| 17 | . Emergent Abilities of Large Language Models. https://arxiv.org/pdf/2206.07682. Type: Paper (arXiv; TMLR). | E1 | ||
| 18 | . Training Verifiers to Solve Math Word Problems. https://arxiv.org/pdf/2110.14168. Type: Paper (arXiv). | E1 | ||
| 19 | . Measuring Mathematical Problem Solving With the MATH Dataset. https://arxiv.org/pdf/2103.03874. Type: Paper (arXiv). | E1 | ||
| 20 | . Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. https://arxiv.org/pdf/2210.09261. Type: Paper (arXiv; ACL Findings). | E1 | ||
| 21 | . Evaluating Large Language Models Trained on Code. https://arxiv.org/pdf/2107.03374. Type: Paper (arXiv). | E1 |
11. Uncertainty log
Overall uncertainty of this entry, bound to the verification log of 2026-08-28 plus TruthfulQA 2026-08-29, GPQA 2026-09-05, Wei 2026-09-06, GSM8K/MATH 2026-09-08, BBH 2026-09-09 and HumanEval 2026-09-12 (source 21). 21 full-text openings. Not used as warrant: aggregators. OpenAI SWE-Verified page: fetch timeout, no warrant.
- Established (layer 1): HELM/Arena/SWE-bench/LCB constructs; TruthfulQA; GPQA (E1); Wei emergence definition (E1) beside Schaeffer; GSM8K/MATH paper measurements (E1); BBH/Suzgun CoT vs answer-only (E1); HumanEval/Chen instrument + Codex pass@k 2021 (E1); MMLU-Pro; contamination; Alzahrani; Arena style control.
- Claimed (layer 2): MMLU-Pro 72.6% GPT-4o; Arena 240k votes Jan 2024; style-control shifts Aug 2024; LCB contamination drops.
- Constrained (layer 3): live leaderboards 2026 not opened (incl. GSM8K/MATH/BBH/HumanEval); SWE-Verified timeout; HumanEval body E1, live scores/contamination 2026 E0; HLE E0; GPQA/GSM8K/MATH bodies E1, live scores E0.
Not opened (not a warrant)
- OpenAI Why we no longer evaluate SWE-bench Verified — timeout; 59.4% E0.
- OpenAI Introducing SWE-bench Verified.
- Aleithan SWE-bench+; Deng SWE-Bench Pro (sister log, E0 here).
- Chen HumanEval opened 2026-09-12 (source 21); Liu HumanEval+ still unopened; Wei Emergent 2022 opened (source 17); GSM8K/MATH opened 2026-09-08 (sources 18–19); BBH opened 2026-09-09 (source 20); BIG-bench full suite still broader than BBH; HLE (GPQA body opened, source 16).
- LMSYS / LMArena live leaderboard 2026; HELM live result tables.
- LiveCodeBench live leaderboard / ICLR 2025.
- GPT-4 Technical Report; Dubois AlpacaEval-LC; Dynabench.