Dossier 12
Linkages
v1.0-draft · unchecked
- E1 confirmed
- E2 single-source
- E3 claimed
- E4 disputed
- E0 unknown
1. As of
- Date
- Version
- v1.0-draft
- Author / model
- Atlas generator
- Reviewer
- unchecked (Josef)
2. In one sentence
Fields couple where the same opened sources carry: more compute lowers training loss and raises datacenter electricity; models fill agent evals far below humans and speed HITL tasks; AlphaFold speeds structure hypotheses without replacing experiment or wet-lab design. The remaining acceleration stories (humanoids, clinic gene therapy, quantum utility, battery UPS as AI grid) are separate or unwarranted in this stock.
Established now · E1 / E2
3. What works today
01 ↔ 02: Compute-optimal training couples models to chips and power
-
E1
Kaplan: cross-entropy loss falls as a power law in N, D and C. Hoffmann/Chinchilla: for a given compute budget, optimal model size and token count scale roughly equally (not 73/27). More compute means more parameters and more data in both papers. Coupling note from the energy log: that is the warrant from dossier 01, not re-measured here.
Kaplan et al. Scaling Laws for Neural Language Models. https://arxiv.org/html/2001.08361. As of 2020-01-23. Checked 2026-08-28. Type: Paper (primary).Hoffmann et al. Training Compute-Optimal Large Language Models. https://arxiv.org/html/2203.15556. As of 2022-03-29. Checked 2026-08-28. Type: Paper. -
E1
Inference energy depends on the task class, not on a per-query average. Luccioni: 88 models, 10 tasks — text classification 0.002 kWh / 1,000; image generation 2.907 kWh / 1,000. Factor >1,000. The same measurement stands in dossiers 01 and 02.
Luccioni, Jernite, Strubell. Power Hungry Processing: Watts Driving the Cost of AI Deployment?. https://arxiv.org/html/2311.16863. As of 2023-11-28. Checked 2026-08-28. Type: Paper (FAccT 2024). -
E1
One training run is measured: BLOOM-176B, 1,082,990 GPU-h, 433,196 kWh dynamic, 50.5 t lifecycle training. Grid intensity beats model size in the paper’s cross-model table.
Luccioni, Viguier, Ligozat. Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. https://arxiv.org/html/2211.02001. As of 2022-11. Checked 2026-08-28. Type: Paper. -
E1
US datacenter electricity: ~60 TWh 2014–16, 76 TWh 2018 (1.9%), 176 TWh 2023 (4.4%). GPU-accelerated servers are the break of the flat era. E1 for 2014–2023; the 2028 range remains a projection. Dossier 02 opens Masanet Science 2020 as of 2026-08-31 for the global 2010–2018 flat phase (figures there) — not the same series as this US break.
Shehabi et al. (LBNL 2024). 2024 United States Data Center Energy Usage Report. https://eta-publications.lbl.gov/sites/default/files/2024-12/lbnl-2024-united-states-data-center-energy-usage-report.pdf. As of 2024-12. Checked 2026-08-28. Type: Official lab report (LBNL-2001637). -
E2
Epoch Trends (as of 2026-02-05): frontier LM training compute 5×/year since 2020; AI chip stock 3.4×/year; GPU FLOP/s/W 1.34×/year since 2008. Training compute grows faster than chip efficiency — the coupling noted in the energy log, E2 (research NGO, inputs not audited). The 5×/year LM-frontier dashboard figure is corroborated as of 2026-08-31 by the May 2024 Epoch report in dossier 01 (figures there). Dossier 02 opened the Epoch chips topic overview on 2026-09-01 (3.3× sold-capacity since 2022, figures there) — not a new source URL here; do not collapse 3.3× / 3.4× / 5.3×.
Epoch Trends. Trends in Artificial Intelligence. https://epoch.ai/trends. As of 2026-02-05. Checked 2026-08-28. Type: Dataset/dashboard. -
E2
DeepSeek-V3 self-report: training 2.788 million H800 GPU-hours. A single-run datapoint that frontier runs buy GPU-hours in the millions — not the world energy balance.
DeepSeek-AI. DeepSeek-V3 Technical Report. https://arxiv.org/html/2412.19437. As of 2024-12-27. Checked 2026-08-28. Type: Tech report.
01 ↔ 03: models fill agent evals; autonomy is not shown
-
E1
ReAct and Toolformer are methods with documented limits, not solved autonomy. Toolformer: cannot chain tools; at most one API call per input. ReAct ALFWorld PaLM-540B best-of-6: 71% vs. Act 45%; WebShop 40.0 vs. human 59.6. Dossier 03 log: dossier 01 is not a warrant — these papers were reopened there.
Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/html/2210.03629. As of 2022-10. Checked 2026-08-28. Type: Paper (ICLR 2023).Schick et al. Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/html/2302.04761. As of 2023-02-09. Checked 2026-08-28. Type: Paper. -
E1
Execution evals, paper snapshots: humans ≫ models. GAIA: humans 92%, GPT-4 + plugins 0% on L3. τ-bench gpt-4o pass^1 61.2% retail / 35.2% airline; pass^8 retail <25%. WebArena: GPT-4 agent 14.41% vs. human 78.24%. OSWorld: humans 72.36%, GPT-4 + accessibility tree 12.24%. Live boards 2026 not opened.
Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper. -
E1
Coding agents: the scaffold moves the score more than “the model alone”. SWE-bench Claude 2 + BM25 1.96%; SWE-agent ACI GPT-4 Turbo full 12.47% / Lite 18.00% vs. shell-only Lite 11.00%. SWE-Pro public Claude Sonnet 4.5 43.6% vs. commercial 17.8%; ablation GPT-5 high 25.9% → 8.40% problem-statement only. Autonomy scores are scaffold scores.
Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10-10. Checked 2026-08-28. Type: Paper.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09-18. Checked 2026-08-28. Type: Paper (Scale AI).
01/03 ↔ 11: HITL tasks speed up; employment statistics do not
-
E1
The same model classes, in supervised settings: Noy/Zhang GPT-3.5 writing −37% time / +0.45 SD / 68% unedited (short tasks). Peng Copilot HTTP server −55.8% time, quality unmeasured, n small. Brynjolfsson: resolutions/hour +13.8%, novices +34%, experts ~0. Human-in-the-loop by design — that is not the autonomy of dossier 03.
Noy & Zhang. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf. As of MIT WP 2023-03-02. Checked 2026-08-28. Type: Working paper.Peng et al. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. https://arxiv.org/html/2302.06590. As of arXiv:2302.06590, experiment May–Jun 2022. Checked 2026-08-28. Type: Paper.Brynjolfsson et al. Generative AI at Work. https://www.nber.org/system/files/working_papers/w31161/w31161.pdf. As of NBER WP 31161, revised Nov 2023. Checked 2026-08-28. Type: Working paper. -
E1
US Census BTOS 2026: firm AI use 17–20% nationally. OECD 2023: little evidence of significant negative employment effects due to AI (evidence without generative AI). Autor 2015: automation substitutes routine tasks and complements non-routine; employment-to-population did not collapse in the 20th century. Coupling: task acceleration is measured; displacement in 2026 is not.
Census BTOS. Large Firms With at Least 20 Employees Biggest AI Users. https://www.census.gov/library/stories/2026/05/ai-use-businesses.html. As of 2026-05-26 (BTOS 14 Dec 2025 – 3 May 2026). Checked 2026-08-28. Type: Official statistics story.OECD EO 2023. OECD Employment Outlook 2023: Artificial Intelligence and the Labour Market. https://www.oecd.org/content/dam/oecd/en/publications/reports/2023/07/oecd-employment-outlook-2023_904bcef3/08785bba-en.pdf. As of 2023-07-11. Checked 2026-08-28. Type: International report.Autor 2015. Why Are There Still So Many Jobs? The History and Future of Workplace Automation. https://economics.mit.edu/sites/default/files/publications/why%20are%20there%20still%20jobs%202014.pdf. As of JEP 29(3), 2015. Checked 2026-08-28. Type: Review/essay (JEP).
06: learned models speed hypotheses, not the wet lab
-
E1
AlphaFold2 on CASP14: median backbone 0.96 Å vs. next-best 2.8 Å. CASP organisers: solution … at least for single proteins. Terwilliger, title and finding: AF predictions are valuable hypotheses and accelerate but do not replace experimental structure determination (map–model 0.56 vs. 0.86). That is the only literally evidenced acceleration between learned models and biology in this stock.
Jumper et al. Highly accurate protein structure prediction with AlphaFold. https://www.nature.com/articles/s41586-021-03819-2. As of 2021-07-15. Checked 2026-08-28. Type: Paper.Kryshtafovych et al. Critical Assessment of Methods of Protein Structure Prediction (CASP) – Round XIV. https://doi.org/10.1002/prot.26237. As of 2021-12. Checked 2026-08-28. Type: Community overview.Terwilliger et al. AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. https://www.nature.com/articles/s41592-023-02087-4. As of 2023-11-30. Checked 2026-08-28. Type: Paper. -
E1
Prediction is not design. AF2 takes a sequence and predicts a structure. RFdiffusion generates sequences; wet-lab binder success 19% (BLI, 95 designs × 5 targets). AF2 in design pipelines is a filter or loss, not the designer. ProteinMPNN opened in dossier 06 as of 2026-08-30 (inverse folding; figures there). ESMFold bioRxiv as of 2026-08-31 (E2); RoseTTAFold as of 2026-09-01 (Baek AAM E2); Nobel Chemistry 2024 official as of 2026-09-02 (E1, figures in D06). AlphaMissense still E0. No new source URLs here.
Jumper et al. Highly accurate protein structure prediction with AlphaFold. https://www.nature.com/articles/s41586-021-03819-2. As of 2021-07-15. Checked 2026-08-28. Type: Paper.Watson et al. De novo design of protein structure and function with RFdiffusion. https://www.nature.com/articles/s41586-023-06415-8.pdf. As of 2023-07-11. Checked 2026-08-28. Type: Paper.
10: progress is progress on an instrument — shared by 01 and 03
-
E1
HELM: accuracy conceals other desiderata; prompting swings 30% → 80%. SWE-bench measures fail-to-pass tests, not software engineering (2023 SOTA 1.96%). Schaeffer: many emergence claims are metric-induced. Dossier 10 measures the instrument; dossiers 01/03 read the same evals as capability. The coupling is methodological, not causal.
Liang et al. Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. As of 2022-11 / TMLR 2023. Checked 2026-08-28. Type: Paper.Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10-10. Checked 2026-08-28. Type: Paper.Schaeffer et al. Are Emergent Abilities of Large Language Models a Mirage?. https://arxiv.org/html/2304.15004. As of 2023-04 / NeurIPS 2023. Checked 2026-08-28. Type: Paper. -
E2
LiveCodeBench: time windows are the hardest contamination control — and only for contest problems (DeepSeek/GPT-4o drop after cutoff). Arena measures preference (240k votes to Jan 2024), not truth. Live ranks 2026 were opened in none of the three logs.
Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper.Chiang et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. As of 2024-03-07 (ICML 2024). Checked 2026-08-28. Type: Paper.
04 ↔ 02: classical GPUs catch NISQ advantage claims
-
E1
Pan, Chen, Zhang: 10⁶ samples of the 53-qubit/20-cycle Sycamore circuit in 15 hours on 512 GPUs, fidelity ≈0.0037. Gao: XEB spoofing in seconds on a GPU (2–12% of experimental XEB values). The same GPU stack that trains models is the classical counter to the supremacy/XEB narrative. That is not AI accelerating quantum computers.
Pan, Chen, Zhang. Solving the Sampling Problem of the Sycamore Quantum Circuits. https://arxiv.org/html/2111.03011. As of 2021-11. Checked 2026-08-28. Type: Paper.Gao et al. Limitations of Linear Cross-Entropy as a Measure for Quantum Advantage. https://arxiv.org/html/2112.01657. As of 2021-12. Checked 2026-08-28. Type: Paper. -
E1
NISQ remains noisy intermediate-scale (Preskill). Willow QEC is below-threshold memory, not a logical algorithm. NIST FIPS 203/204/205 exist because of a future CRQC threat, not because RSA is broken. Dossier 04 opens the FIPS 203/204/205 bodies as of 2026-09-02/03 (ML-KEM/ML-DSA/SLH-DSA sizes) — figures only there; no new source URL here.
Preskill. Quantum Computing in the NISQ era and beyond. https://arxiv.org/html/1801.00862. As of 2018. Checked 2026-08-28. Type: Review (Quantum 2, 79).Acharya et al. Quantum error correction below the surface code threshold. https://arxiv.org/html/2408.13687. As of 2024-12-09. Checked 2026-08-28. Type: Paper (Nature 638).NIST PQC. Announcing Approval of Three Federal Information Processing Standards (FIPS) for Post-Quantum Cryptography. https://www.nist.gov/news-events/news/2024/08/announcing-approval-three-federal-information-processing-standards-fips. As of 2024-08-13. Checked 2026-08-28. Type: Agency.
Claimed · E3
4. What is claimed, not shown
IEA: agents and reasoning as expensive queries; HBM and capex
-
E2
IEA base/update: global datacenter electricity ~415 TWh (2024) or 485 TWh (2025) → ~950 TWh (2030), ~3% of world electricity. AI-focused DCs triple; +50% in 2025 vs. +17% all DCs. 2026 figures per Key Questions: estimates. Methods annex not fully read. LBNL 2025: US 2024 192 TWh = 4.7%; 2030 reference 649 TWh — 2030 levels not harmonised with IEA.
IEA. Energy and AI — Executive summary. https://www.iea.org/reports/energy-and-ai/executive-summary. As of 2025-04. Checked 2026-08-28. Type: Official report.IEA Key Questions. Key Questions on Energy and AI — Executive Summary. https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary. As of 2026. Checked 2026-08-28. Type: Official report (exec summary).Smith et al. (LBNL 2025). United States Data Center Energy Usage Report: 2025 Update. https://doi.org/10.71468/P1RP4F. As of 2026-06-18. Checked 2026-08-28. Type: Official lab report (LBNL-2001758). -
E2
IEA Key Questions: energy per task has fallen “at least an order of magnitude annually”. Video/reasoning/agents: “hundreds or thousands of times” more energy per query. E2 for direction and Luccioni task classes; E3 for the exact order of magnitude/year while methods remain unread. That is the claimed coupling agents → power, not measured autonomy.
IEA Key Questions. Key Questions on Energy and AI — Executive Summary. https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary. As of 2026. Checked 2026-08-28. Type: Official report (exec summary).Luccioni, Jernite, Strubell. Power Hungry Processing: Watts Driving the Cost of AI Deployment?. https://arxiv.org/html/2311.16863. As of 2023-11-28. Checked 2026-08-28. Type: Paper (FAccT 2024). -
E2
HBM shortage “over the past six months”, expected until at least end-2027. Capex of the largest tech firms > USD 400 billion in 2025, +75% expected 2026. Grid: ~20% of planned DC projects at delay risk. Satellite method not independently replicated.
IEA Key Questions. Key Questions on Energy and AI — Executive Summary. https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary. As of 2026. Checked 2026-08-28. Type: Official report (exec summary).IEA Security. AI and energy security. https://www.iea.org/reports/energy-and-ai/ai-and-energy-security. As of 2025-04. Checked 2026-08-28. Type: Official report (chapter).
Product autonomy and humanoid demos
-
E3
Devin: “the first AI software engineer” (E3 marketing) against 13.86% on a 25% subset of SWE-bench (E2 vendor measurement), in the same range as SWE-agent 12.47% full.
Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report. -
E2
Anthropic computer use: itself “experimental—at times cumbersome and error-prone.” OpenAI CUA: “we don’t expect CUA to perform reliably in all scenarios just yet”; confirmation and Watch mode. ChatGPT agent: permission before consequential actions; refuse bank transfers; “still in its early stages”. Product autonomy is human-in-the-loop in the vendors’ own texts.
Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch. -
E3
Figure F.02 “contributed to 30,000+ X3” and Boston Dynamics Atlas CES 2026 (56 DoF, prototype on stage / currently in development) are vendor demos. IFR infographic: humanoids will not compete with industrial robots on speed, precision, reliability. Foundation models in the physical world are only a pointer in the robotics log — π0/RT-X not opened.
Figure F.02. F.02 Contributed to the Production of 30,000 Cars at BMW. https://www.figure.ai/news/production-at-bmw. As of 2025-11-19. Checked 2026-08-28. Type: Company report.Atlas CES. Boston Dynamics Unveils New Atlas Robot to Revolutionize Industry. https://bostondynamics.com/blog/boston-dynamics-unveils-new-atlas-robot-to-revolutionize-industry/. As of 2026-01-05. Checked 2026-08-28. Type: Company report.IFR humanoid info. Humanoid Robots — Vision and Reality (Infographic). https://ifr.org/downloads/press_docs/Humanoids_Position_Infograph_2025.pdf. As of 2025-08. Checked 2026-08-28. Type: Association PDF.
Exposure scores and QC resources — scenario, not the present
-
E2
Eloundou: ~80% of the US workforce with ≥10% of tasks affected; “We do not make predictions about the development or adoption timeline.” Frey 2013: 47% high-risk “over some unspecified number of years”. Both are occupation scores, not a 2026 displacement measurement. Eloundou: ζ 46% needs software co-invention.
Eloundou et al. GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. https://arxiv.org/html/2303.10130. As of arXiv:2303.10130 (rev. Aug 2023). Checked 2026-08-28. Type: Paper (OpenAI).Frey & Osborne. The Future of Employment: How Susceptible are Jobs to Computerisation?. https://oms-www.files.svdcdn.com/production/downloads/academic/The_Future_of_Employment.pdf. As of Oxford Martin WP 2013-09-17. Checked 2026-08-28. Type: Working paper. -
E3
Gidney 2025: RSA-2048 with <1 million noisy qubits in <1 week under assumptions 0.1% gate error, 1 µs surface-code cycle. Estimate, not a device. Opened chips in dossier 04 (53 / 105 / 127 / 30 ions / 280 atoms) sit 3–4 orders of magnitude below.
Gidney 2025. How to factor 2048 bit RSA integers with less than a million noisy qubits. https://arxiv.org/html/2505.15917. As of 2025-05. Checked 2026-08-28. Type: Paper (resource estimate).
Constrained · Limit
5. Bottleneck and limit
Jevons and the chip wall: efficiency rises, net power does too
-
E2
Efficiency/task rises (IEA; Luccioni task classes; Epoch FLOP/s/W). At the same time, more expensive tasks and more chips (3.4×/year stock). IEA central path: DC electricity doubles by 2030. LBNL: the US era of flat kWh is over. Causal share of AI vs. rest-of-cloud vs. crypto is not finely resolved. Patterson “plateau then shrink” (Google internal, E2) stands against this net rise and is not a 2026 world forecast (E3).
IEA Key Questions. Key Questions on Energy and AI — Executive Summary. https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary. As of 2026. Checked 2026-08-28. Type: Official report (exec summary).Luccioni, Jernite, Strubell. Power Hungry Processing: Watts Driving the Cost of AI Deployment?. https://arxiv.org/html/2311.16863. As of 2023-11-28. Checked 2026-08-28. Type: Paper (FAccT 2024).Epoch Trends. Trends in Artificial Intelligence. https://epoch.ai/trends. As of 2026-02-05. Checked 2026-08-28. Type: Dataset/dashboard.Shehabi et al. (LBNL 2024). 2024 United States Data Center Energy Usage Report. https://eta-publications.lbl.gov/sites/default/files/2024-12/lbnl-2024-united-states-data-center-energy-usage-report.pdf. As of 2024-12. Checked 2026-08-28. Type: Official lab report (LBNL-2001637).Patterson et al. The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink. https://arxiv.org/abs/2204.05149. As of 2022-04-11. Checked 2026-08-28. Type: Paper (Google/Berkeley). -
E1
Memory wall: compute 2× / 2.3 years vs. bandwidth 2× / ~4 years (Epoch 2023). H100 max TDP up to 700 W (spec, not field utilization). Dossier 02: TSMC is producing 5.5-reticle CoWoS now (2026 Symposium PR); kwpm still E0. ASML 48 EUV in 2025. The chip limit of models sits in packaging, HBM, lithography — not only GPU TDP. IEA HBM until ≥2027 remains E2. No new source URL here.
Hobbhahn, Heim, Aydos. Trends in machine learning hardware. https://epoch.ai/publications/trends-in-machine-learning-hardware. As of 2023-11-09. Checked 2026-08-28. Type: Dataset/report.NVIDIA. NVIDIA H100 GPU (product specifications). https://www.nvidia.com/en-us/data-center/h100/. As of 2026-08-28. Checked 2026-08-28. Type: Manufacturer specification.TSMC. 2025 Annual Report, Letter to Shareholders (ch. 1 PDF). https://investor.tsmc.com/static/annualReports/2025/english/pdf/2025_tsmc_ar_e_ch1.pdf. As of 2025. Checked 2026-08-28. Type: Company report.ASML PR. ASML reports €32.7 billion total net sales and €9.6 billion net income in 2025. https://www.asml.com/en/news/press-releases/2026/q4-2025-financial-results. As of 2026-01-28. Checked 2026-08-28. Type: Company report.
Autonomy illusion, scaffold, contamination
-
E2
“Autonomous” in products is not unsupervised by design: confirmation, Watch mode, decline banking. The product model is human-in-the-loop — the same structure as the productivity RCTs in dossier 11, not unsupervised execution on GAIA L3.
Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch. -
E2
OpenAI 2026-02-23: SWE-bench Verified no longer measures frontier coding (contaminated gold patches, 59.4% material test issues on 138 unsolved). SWE-Pro commercial ≪ public. LiveCodeBench: drop after cutoff. High coding scores without these constraints are unreadable. Dossier 10 did not reopen Verified/Pro in its run (E0 there) — the figures here come from 01/03.
OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09-18. Checked 2026-08-28. Type: Paper (Scale AI).Jain et al. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. As of 2024-03 (HTML: 511 problems May 2023–May 2024). Checked 2026-08-28. Type: Paper.
Exposure is not displacement; education: see dossier 11
-
E1
Pattern through BTOS 2026: writing/code/customer service speeds executable tasks in RCT/field. Employment statistics and firm adoption (17–20%) stay behind Frey 47% and Eloundou 80%. Frey 2013+13 years = 2026, and OECD/Census show no 47% event (E4 as observed displacement).
Noy & Zhang. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf. As of MIT WP 2023-03-02. Checked 2026-08-28. Type: Working paper.Peng et al. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. https://arxiv.org/html/2302.06590. As of arXiv:2302.06590, experiment May–Jun 2022. Checked 2026-08-28. Type: Paper.Brynjolfsson et al. Generative AI at Work. https://www.nber.org/system/files/working_papers/w31161/w31161.pdf. As of NBER WP 31161, revised Nov 2023. Checked 2026-08-28. Type: Working paper.Census BTOS. Large Firms With at Least 20 Employees Biggest AI Users. https://www.census.gov/library/stories/2026/05/ai-use-businesses.html. As of 2026-05-26 (BTOS 14 Dec 2025 – 3 May 2026). Checked 2026-08-28. Type: Official statistics story.OECD EO 2023. OECD Employment Outlook 2023: Artificial Intelligence and the Labour Market. https://www.oecd.org/content/dam/oecd/en/publications/reports/2023/07/oecd-employment-outlook-2023_904bcef3/08785bba-en.pdf. As of 2023-07-11. Checked 2026-08-28. Type: International report.Frey & Osborne. The Future of Employment: How Susceptible are Jobs to Computerisation?. https://oms-www.files.svdcdn.com/production/downloads/academic/The_Future_of_Employment.pdf. As of Oxford Martin WP 2013-09-17. Checked 2026-08-28. Type: Working paper. -
E0
Education: see dossier 11, Bastani opened 2026-08-30 (field RCT, exam −17% GPT Base / Tutor ≈ control); PISA/UNESCO still unopened. This synthesis dossier does not copy new source URLs. Productivity RCTs measure output with a tool, not unaided learning.
Prediction is not design and not the clinic
-
E1
Synbio log: protein design is not chassis and not clinic; pointer only, do not write the other dossier. CASGEVY is ex-vivo HSCT after busulfan myeloablation (label 07/2026), not a pill designed by AF. IGSC: no country requires providers to screen. No opened source connects RFdiffusion binders or AF2 to CASGEVY, syn3.0, or DNA screening.
Jumper et al. Highly accurate protein structure prediction with AlphaFold. https://www.nature.com/articles/s41586-021-03819-2. As of 2021-07-15. Checked 2026-08-28. Type: Paper.Watson et al. De novo design of protein structure and function with RFdiffusion. https://www.nature.com/articles/s41586-023-06415-8.pdf. As of 2023-07-11. Checked 2026-08-28. Type: Paper.FDA 2023-12-08. FDA Approves First Gene Therapies to Treat Patients with Sickle Cell Disease. https://content.govdelivery.com/accounts/USFDA/bulletins/37efce0. As of 2023-12-08. Checked 2026-08-28. Type: Agency notice.CASGEVY USPI. CASGEVY (exagamglogene autotemcel) Prescribing Information. https://dailymed.nlm.nih.gov/dailymed/drugInfo.cfm?setid=7c3e12ad-e2fe-4d3f-a630-ea7364d9e846. As of Revised 07/2026. Checked 2026-08-28. Type: Prescribing information.IGSC v3.1. Harmonized Screening Protocol Version 3.1. https://genesynthesisconsortium.org/wp-content/uploads/IGSC-Harmonized-Screening-Protocol-V3.1.pdf. As of 2026-06-01. Checked 2026-08-28. Type: Industry protocol. -
E0
Dossier 06 opens RoseTTAFold as of 2026-09-01 (Baek Science 2021 AAM) and the Nobel Chemistry 2024 press release as of 2026-09-02 (Baker design / Hassabis+Jumper prediction, E1). Figures only there. No new source URL here. AlphaMissense still E0.
Jumper et al. Highly accurate protein structure prediction with AlphaFold. https://www.nature.com/articles/s41586-021-03819-2. As of 2021-07-15. Checked 2026-08-28. Type: Paper.
Embodiment is not agent; humanoids not in the stock
-
E1
What ships: 542,076 industrial robot installations in 2024, stock 4.66 million (IFR). OpenAI Rubik: full scramble 20%, vision-only 0%, N=10, cage. Reality gap remains the central transfer diagnosis. Robotics log: foundation models / agents in the physical world are embodiment, not SWE-bench. Teleoperation is not a robot under ISO 8373.
IFR WR 2025 Exec. World Robotics 2025 — Industrial Robots Executive Summary. https://ifr.org/img/worldrobotics/Executive_Summary_WR_2025_Industrial_Robots.pdf. As of 2025. Checked 2026-08-28. Type: Association PDF.OpenAI Rubik. Solving Rubik's Cube with a Robot Hand. https://arxiv.org/html/1910.07113. As of 2019-10. Checked 2026-08-28. Type: Paper.Reality Gap. The Reality Gap in Robotics: Challenges, Solutions, and Best Practices. https://arxiv.org/html/2510.20808. As of 2025-10. Checked 2026-08-28. Type: Paper (survey). -
E2
IFR: humanoids “no massive use today”. Infographic: will not compete on speed/precision/reliability; battery ≠ a full workday; ISO safety for legged robots “just started”. Tesla Optimus: access denied = E0 in dossier 08. π0/RT-X not opened.
IFR humanoid info. Humanoid Robots — Vision and Reality (Infographic). https://ifr.org/downloads/press_docs/Humanoids_Position_Infograph_2025.pdf. As of 2025-08. Checked 2026-08-28. Type: Association PDF.IFR humanoid press. Humanoid Robots: Vision and Reality Paper Published by IFR. https://ifr.org/ifr-press-releases/news/humanoid-robots-vision-and-reality-paper-published-by-ifr. As of 2025-08-14. Checked 2026-08-28. Type: Association press release.
Quantum computing is not quantum sensing; battery UPS is not the AI grid
-
E1
Degen: type I/II often near applications; type III (entanglement beyond the classical limit) is the strict definition and not the deployed class. Cs defines the SI second; GPS payload is rubidium microwave. NISQ computers are experiments without an FT algorithm. No opened source shows sensing accelerating computers or vice versa — the logs explicitly keep the stacks apart.
Degen, Reinhard, Cappellaro. Quantum sensing. https://arxiv.org/html/1611.02427. As of 2017-06. Checked 2026-08-28. Type: Review (RMP 89).NIST How Do We Know. How Do We Know What Time It Is?. https://www.nist.gov/atomic-clocks/how-do-we-know-what-time-it. As of 2025-05-28. Checked 2026-08-28. Type: Agency.Preskill. Quantum Computing in the NISQ era and beyond. https://arxiv.org/html/1801.00862. As of 2018. Checked 2026-08-28. Type: Review (Quantum 2, 79). -
E1
Batteries log, coupling energy/chips: UPS additions 45 GW in 2025 are short-time backup, not grid storage. IEA: batteries good for 1–8 h; deployments ~2 h. Li-ion ships at TWh scale (1 TWh 2024, EV+storage), China 80–85% of cells. That does not solve the datacenter power path. ML-designed cells: no opened source.
IEA GEO 2025. Electric vehicle batteries — Global EV Outlook 2025. https://www.iea.org/reports/global-ev-outlook-2025/electric-vehicle-batteries. As of 2025. Checked 2026-08-28. Type: Official report.IEA GER 2026. Technology: Battery storage — Global Energy Review 2026. https://www.iea.org/reports/global-energy-review-2026/technology-battery-storage. As of 2026. Checked 2026-08-28. Type: Official report.
6. Actors and incentives
Who measures the coupling, who sells it
-
E1
LBNL measures US DC electricity bottom-up. TSMC reports wafer mix; ASML EUV units. IFR counts industrial robots under ISO 8373. These are the physical counters, not the model leaderboards.
Shehabi et al. (LBNL 2024). 2024 United States Data Center Energy Usage Report. https://eta-publications.lbl.gov/sites/default/files/2024-12/lbnl-2024-united-states-data-center-energy-usage-report.pdf. As of 2024-12. Checked 2026-08-28. Type: Official lab report (LBNL-2001637).TSMC. 2025 Annual Report, Letter to Shareholders (ch. 1 PDF). https://investor.tsmc.com/static/annualReports/2025/english/pdf/2025_tsmc_ar_e_ch1.pdf. As of 2025. Checked 2026-08-28. Type: Company report.ASML PR. ASML reports €32.7 billion total net sales and €9.6 billion net income in 2025. https://www.asml.com/en/news/press-releases/2026/q4-2025-financial-results. As of 2026-01-28. Checked 2026-08-28. Type: Company report.IFR WR 2025 Exec. World Robotics 2025 — Industrial Robots Executive Summary. https://ifr.org/img/worldrobotics/Executive_Summary_WR_2025_Industrial_Robots.pdf. As of 2025. Checked 2026-08-28. Type: Association PDF. -
E2
IEA projects DC electricity and HBM. Epoch curates training compute and chip stock. OpenAI withdraws SWE-Verified as a frontier measure; Scale AI publishes SWE-Pro with a public/commercial split. Incentive: whoever builds the instrument defines what progress means.
IEA. Energy and AI — Executive summary. https://www.iea.org/reports/energy-and-ai/executive-summary. As of 2025-04. Checked 2026-08-28. Type: Official report.Epoch Trends. Trends in Artificial Intelligence. https://epoch.ai/trends. As of 2026-02-05. Checked 2026-08-28. Type: Dataset/dashboard.OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09-18. Checked 2026-08-28. Type: Paper (Scale AI). -
E1
DeepMind ships AF2; Baker lab measures RFdiffusion wet lab. Noy/Zhang, Peng (Microsoft/GitHub), Brynjolfsson/Li/Raymond measure HITL productivity. CASP scores AF independently.
Jumper et al. Highly accurate protein structure prediction with AlphaFold. https://www.nature.com/articles/s41586-021-03819-2. As of 2021-07-15. Checked 2026-08-28. Type: Paper.Watson et al. De novo design of protein structure and function with RFdiffusion. https://www.nature.com/articles/s41586-023-06415-8.pdf. As of 2023-07-11. Checked 2026-08-28. Type: Paper.Noy & Zhang. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf. As of MIT WP 2023-03-02. Checked 2026-08-28. Type: Working paper.Peng et al. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. https://arxiv.org/html/2302.06590. As of arXiv:2302.06590, experiment May–Jun 2022. Checked 2026-08-28. Type: Paper.Brynjolfsson et al. Generative AI at Work. https://www.nber.org/system/files/working_papers/w31161/w31161.pdf. As of NBER WP 31161, revised Nov 2023. Checked 2026-08-28. Type: Working paper. -
E2
Cognition, Anthropic, OpenAI sell autonomy on pages that carry caveats. Figure and Boston Dynamics are demo sources; their caveats sit on the same pages.
Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.Figure F.02. F.02 Contributed to the Production of 30,000 Cars at BMW. https://www.figure.ai/news/production-at-bmw. As of 2025-11-19. Checked 2026-08-28. Type: Company report.Atlas CES. Boston Dynamics Unveils New Atlas Robot to Revolutionize Industry. https://bostondynamics.com/blog/boston-dynamics-unveils-new-atlas-robot-to-revolutionize-industry/. As of 2026-01-05. Checked 2026-08-28. Type: Company report.
7. State of the dispute
What is collapsed into one and does not coincide in the sources
-
E1
Kaplan vs. Chinchilla: the power law of training loss stands in both; the compute-optimal N/D allocation is revised. Both full texts opened in dossier 01.
Kaplan et al. Scaling Laws for Neural Language Models. https://arxiv.org/html/2001.08361. As of 2020-01-23. Checked 2026-08-28. Type: Paper (primary).Hoffmann et al. Training Compute-Optimal Large Language Models. https://arxiv.org/html/2203.15556. As of 2022-03-29. Checked 2026-08-28. Type: Paper. -
E2
Patterson “plateau then shrink” (Google internal, 4Ms) against the IEA/LBNL net rise. E2 internal; E3 as a 2026 world forecast.
Patterson et al. The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink. https://arxiv.org/abs/2204.05149. As of 2022-04-11. Checked 2026-08-28. Type: Paper (Google/Berkeley).IEA. Energy and AI — Executive summary. https://www.iea.org/reports/energy-and-ai/executive-summary. As of 2025-04. Checked 2026-08-28. Type: Official report.Shehabi et al. (LBNL 2024). 2024 United States Data Center Energy Usage Report. https://eta-publications.lbl.gov/sites/default/files/2024-12/lbnl-2024-united-states-data-center-energy-usage-report.pdf. As of 2024-12. Checked 2026-08-28. Type: Official lab report (LBNL-2001637). -
E2
Product claims of agent autonomy (Devin E3) stand against GAIA, τ-bench and SWE-Pro (measured gaps, E1/E2). Dossier 01 and 03 logs run the same dispute.
Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report.Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09-18. Checked 2026-08-28. Type: Paper (Scale AI). -
E4
Frey 47% “at risk” (2013, technical feasibility) vs. OECD 2023 “no signs of slowing labour demand (yet)” vs. Census 2026 adoption 17–20%. Autor 2015 cites exactly Frey as overstated ML optimism.
Frey & Osborne. The Future of Employment: How Susceptible are Jobs to Computerisation?. https://oms-www.files.svdcdn.com/production/downloads/academic/The_Future_of_Employment.pdf. As of Oxford Martin WP 2013-09-17. Checked 2026-08-28. Type: Working paper.OECD EO 2023. OECD Employment Outlook 2023: Artificial Intelligence and the Labour Market. https://www.oecd.org/content/dam/oecd/en/publications/reports/2023/07/oecd-employment-outlook-2023_904bcef3/08785bba-en.pdf. As of 2023-07-11. Checked 2026-08-28. Type: International report.Census BTOS. Large Firms With at Least 20 Employees Biggest AI Users. https://www.census.gov/library/stories/2026/05/ai-use-businesses.html. As of 2026-05-26 (BTOS 14 Dec 2025 – 3 May 2026). Checked 2026-08-28. Type: Official statistics story.Autor 2015. Why Are There Still So Many Jobs? The History and Future of Workplace Automation. https://economics.mit.edu/sites/default/files/publications/why%20are%20there%20still%20jobs%202014.pdf. As of JEP 29(3), 2015. Checked 2026-08-28. Type: Review/essay (JEP). -
E1
Sycamore experiment real, 10,000-year claim not. Willow QEC = memory, not algorithm. Quantum-sensing log: do not collapse Cs time, optical lab clocks (including Bothwell millimetre redshift in D05 as of 2026-09-03), GPS Rb, SQUID vs. OPM, gravimeter vs. satellite into “quantum sensing” — and do not collapse them with NISQ computers. Figures only in D05; no new source URL here.
Pan, Chen, Zhang. Solving the Sampling Problem of the Sycamore Quantum Circuits. https://arxiv.org/html/2111.03011. As of 2021-11. Checked 2026-08-28. Type: Paper.Acharya et al. Quantum error correction below the surface code threshold. https://arxiv.org/html/2408.13687. As of 2024-12-09. Checked 2026-08-28. Type: Paper (Nature 638).Degen, Reinhard, Cappellaro. Quantum sensing. https://arxiv.org/html/1611.02427. As of 2017-06. Checked 2026-08-28. Type: Review (RMP 89).
8. Open questions
- ESMFold Science/PMC body (bioRxiv opened E2 in D06 as of 2026-08-31; RoseTTAFold AAM as of 2026-09-01; Nobel press as of 2026-09-02). AlphaMissense still E0.
- π0 / RT-X / Tesla Optimus primary pages — without them foundation-model embodiment stays E0/E3-vendor.
- PISA/UNESCO/OECD Education — Bastani opened in dossier 11; otherwise the education half of the title stays empty.
- IEA Energy and AI methods annex — how 415/945 and “order of magnitude/year” are built.
- Epoch chips-topic-overview opened in D02 2026-09-01; CoWoS 5.5 in D02 as of 2026-09-02 — figures there. Compute price series still a gap.
- Quantum ML / ML-for-QEC primary sources — not opened in dossier 04.
- Janek & Zeier 2023 full text — ML-designed solid-state cells not claimed, only the gap.
- Live leaderboards 2026 (GAIA HF, SWE-Pro public, Arena, HELM) — screenshot plus date.
- in-vivo gene-therapy primary sources — E0 in dossier 09; do not couple with AF design.
- WebVoyager original; agent system-card PDFs (Operator, ChatGPT agent).
9. Changes
- v1.0-draft2026-09-03: light coupling: D04 FIPS 204/205 bodies (ML-DSA/SLH-DSA sizes), D05 Bothwell millimetre redshift — figures in the sister dossiers, no new source URLs here.
- v1.0-draft2026-09-02: light coupling: D04 FIPS 203 body (ML-KEM sizes), D02 CoWoS 5.5-reticle, D06 Nobel Chemistry 2024 official — figures in the sister dossiers, no new source URLs here. Cleared stale RoseTTAFold-E0 line; AlphaMissense still E0.
- v1.0-draft2026-09-01: light coupling: D02 Epoch chips overview (3.3× ≠ 5.3×), D06 RoseTTAFold — figures in the sister dossiers, no new source URLs here.
- v1.0-draft2026-08-31: light coupling sentences for D01 Epoch report, D02 Masanet, D06 ESMFold — figures in the sister dossiers, no new sources here.
- v1.0-draftFirst version as a synthesis of dossiers 01–11 and their 2026-08-28 verification logs. No new sources opened. Couplings only where sister logs carry the same warrant or explicitly note a pointer.
10. Sources
| No. | Source | As of | Checked | Grade |
|---|---|---|---|---|
| 1 | . Scaling Laws for Neural Language Models. https://arxiv.org/html/2001.08361. Type: Paper (primary). | E1 | ||
| 2 | . Training Compute-Optimal Large Language Models. https://arxiv.org/html/2203.15556. Type: Paper. | E1 | ||
| 3 | . Power Hungry Processing: Watts Driving the Cost of AI Deployment?. https://arxiv.org/html/2311.16863. Type: Paper (FAccT 2024). | E1 | ||
| 4 | . Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. https://arxiv.org/html/2211.02001. Type: Paper. | E1 | ||
| 5 | . Trends in Artificial Intelligence. https://epoch.ai/trends. Type: Dataset/dashboard. | E2 | ||
| 6 | . Trends in machine learning hardware. https://epoch.ai/publications/trends-in-machine-learning-hardware. Type: Dataset/report. | E2 | ||
| 7 | . 2024 United States Data Center Energy Usage Report. https://eta-publications.lbl.gov/sites/default/files/2024-12/lbnl-2024-united-states-data-center-energy-usage-report.pdf. Type: Official lab report (LBNL-2001637). | E1 | ||
| 8 | . United States Data Center Energy Usage Report: 2025 Update. https://doi.org/10.71468/P1RP4F. Type: Official lab report (LBNL-2001758). | E2 | ||
| 9 | . Energy and AI — Executive summary. https://www.iea.org/reports/energy-and-ai/executive-summary. Type: Official report. | E2 | ||
| 10 | . Key Questions on Energy and AI — Executive Summary. https://www.iea.org/reports/key-questions-on-energy-and-ai/executive-summary. Type: Official report (exec summary). | E2 | ||
| 11 | . AI and energy security. https://www.iea.org/reports/energy-and-ai/ai-and-energy-security. Type: Official report (chapter). | E2 | ||
| 12 | . NVIDIA H100 GPU (product specifications). https://www.nvidia.com/en-us/data-center/h100/. Type: Manufacturer specification. | E1 | ||
| 13 | . 2025 Annual Report, Letter to Shareholders (ch. 1 PDF). https://investor.tsmc.com/static/annualReports/2025/english/pdf/2025_tsmc_ar_e_ch1.pdf. Type: Company report. | E1 | ||
| 14 | . ASML reports €32.7 billion total net sales and €9.6 billion net income in 2025. https://www.asml.com/en/news/press-releases/2026/q4-2025-financial-results. Type: Company report. | E1 | ||
| 15 | . The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink. https://arxiv.org/abs/2204.05149. Type: Paper (Google/Berkeley). | E2 | ||
| 16 | . DeepSeek-V3 Technical Report. https://arxiv.org/html/2412.19437. Type: Tech report. | E2 | ||
| 17 | . ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/html/2210.03629. Type: Paper (ICLR 2023). | E1 | ||
| 18 | . Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/html/2302.04761. Type: Paper. | E1 | ||
| 19 | . GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. Type: Paper. | E1 | ||
| 20 | . τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. Type: Paper. | E1 | ||
| 21 | . SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. Type: Paper. | E1 | ||
| 22 | . SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. Type: Paper. | E1 | ||
| 23 | . SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. Type: Paper (Scale AI). | E2 | ||
| 24 | . OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. Type: Paper. | E1 | ||
| 25 | . WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. Type: Paper. | E1 | ||
| 26 | . Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. Type: Company report. | E2 | ||
| 27 | . Holistic Evaluation of Language Models. https://arxiv.org/html/2211.09110. Type: Paper. | E1 | ||
| 28 | . LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. https://arxiv.org/html/2403.07974. Type: Paper. | E2 | ||
| 29 | . Are Emergent Abilities of Large Language Models a Mirage?. https://arxiv.org/html/2304.15004. Type: Paper. | E1 | ||
| 30 | . Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. https://arxiv.org/html/2403.04132. Type: Paper. | E2 | ||
| 31 | . Highly accurate protein structure prediction with AlphaFold. https://www.nature.com/articles/s41586-021-03819-2. Type: Paper. | E1 | ||
| 32 | . AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. https://www.nature.com/articles/s41592-023-02087-4. Type: Paper. | E1 | ||
| 33 | . De novo design of protein structure and function with RFdiffusion. https://www.nature.com/articles/s41586-023-06415-8.pdf. Type: Paper. | E1 | ||
| 34 | . Critical Assessment of Methods of Protein Structure Prediction (CASP) – Round XIV. https://doi.org/10.1002/prot.26237. Type: Community overview. | E1 | ||
| 35 | . Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence. https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf. Type: Working paper. | E1 | ||
| 36 | . The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. https://arxiv.org/html/2302.06590. Type: Paper. | E1 | ||
| 37 | . Generative AI at Work. https://www.nber.org/system/files/working_papers/w31161/w31161.pdf. Type: Working paper. | E1 | ||
| 38 | . Large Firms With at Least 20 Employees Biggest AI Users. https://www.census.gov/library/stories/2026/05/ai-use-businesses.html. Type: Official statistics story. | E1 | ||
| 39 | . GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. https://arxiv.org/html/2303.10130. Type: Paper (OpenAI). | E2 | ||
| 40 | . The Future of Employment: How Susceptible are Jobs to Computerisation?. https://oms-www.files.svdcdn.com/production/downloads/academic/The_Future_of_Employment.pdf. Type: Working paper. | E2 | ||
| 41 | . OECD Employment Outlook 2023: Artificial Intelligence and the Labour Market. https://www.oecd.org/content/dam/oecd/en/publications/reports/2023/07/oecd-employment-outlook-2023_904bcef3/08785bba-en.pdf. Type: International report. | E1 | ||
| 42 | . Why Are There Still So Many Jobs? The History and Future of Workplace Automation. https://economics.mit.edu/sites/default/files/publications/why%20are%20there%20still%20jobs%202014.pdf. Type: Review/essay (JEP). | E1 | ||
| 43 | . Solving the Sampling Problem of the Sycamore Quantum Circuits. https://arxiv.org/html/2111.03011. Type: Paper. | E1 | ||
| 44 | . Limitations of Linear Cross-Entropy as a Measure for Quantum Advantage. https://arxiv.org/html/2112.01657. Type: Paper. | E1 | ||
| 45 | . Quantum Computing in the NISQ era and beyond. https://arxiv.org/html/1801.00862. Type: Review (Quantum 2, 79). | E1 | ||
| 46 | . Announcing Approval of Three Federal Information Processing Standards (FIPS) for Post-Quantum Cryptography. https://www.nist.gov/news-events/news/2024/08/announcing-approval-three-federal-information-processing-standards-fips. Type: Agency. | E1 | ||
| 47 | . Quantum error correction below the surface code threshold. https://arxiv.org/html/2408.13687. Type: Paper (Nature 638). | E1 | ||
| 48 | . How to factor 2048 bit RSA integers with less than a million noisy qubits. https://arxiv.org/html/2505.15917. Type: Paper (resource estimate). | E3 | ||
| 49 | . Quantum sensing. https://arxiv.org/html/1611.02427. Type: Review (RMP 89). | E1 | ||
| 50 | . How Do We Know What Time It Is?. https://www.nist.gov/atomic-clocks/how-do-we-know-what-time-it. Type: Agency. | E1 | ||
| 51 | . Electric vehicle batteries — Global EV Outlook 2025. https://www.iea.org/reports/global-ev-outlook-2025/electric-vehicle-batteries. Type: Official report. | E1 | ||
| 52 | . Technology: Battery storage — Global Energy Review 2026. https://www.iea.org/reports/global-energy-review-2026/technology-battery-storage. Type: Official report. | E2 | ||
| 53 | . World Robotics 2025 — Industrial Robots Executive Summary. https://ifr.org/img/worldrobotics/Executive_Summary_WR_2025_Industrial_Robots.pdf. Type: Association PDF. | E1 | ||
| 54 | . Humanoid Robots — Vision and Reality (Infographic). https://ifr.org/downloads/press_docs/Humanoids_Position_Infograph_2025.pdf. Type: Association PDF. | E2 | ||
| 55 | . Humanoid Robots: Vision and Reality Paper Published by IFR. https://ifr.org/ifr-press-releases/news/humanoid-robots-vision-and-reality-paper-published-by-ifr. Type: Association press release. | E2 | ||
| 56 | . Solving Rubik's Cube with a Robot Hand. https://arxiv.org/html/1910.07113. Type: Paper. | E1 | ||
| 57 | . The Reality Gap in Robotics: Challenges, Solutions, and Best Practices. https://arxiv.org/html/2510.20808. Type: Paper (survey). | E1 | ||
| 58 | . F.02 Contributed to the Production of 30,000 Cars at BMW. https://www.figure.ai/news/production-at-bmw. Type: Company report. | E3 | ||
| 59 | . Boston Dynamics Unveils New Atlas Robot to Revolutionize Industry. https://bostondynamics.com/blog/boston-dynamics-unveils-new-atlas-robot-to-revolutionize-industry/. Type: Company report. | E3 | ||
| 60 | . FDA Approves First Gene Therapies to Treat Patients with Sickle Cell Disease. https://content.govdelivery.com/accounts/USFDA/bulletins/37efce0. Type: Agency notice. | E1 | ||
| 61 | . CASGEVY (exagamglogene autotemcel) Prescribing Information. https://dailymed.nlm.nih.gov/dailymed/drugInfo.cfm?setid=7c3e12ad-e2fe-4d3f-a630-ea7364d9e846. Type: Prescribing information. | E1 | ||
| 62 | . Harmonized Screening Protocol Version 3.1. https://genesynthesisconsortium.org/wp-content/uploads/IGSC-Harmonized-Screening-Protocol-V3.1.pdf. Type: Industry protocol. | E1 | ||
| 63 | . Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. Type: Company news. | E2 | ||
| 64 | . Computer-Using Agent. https://openai.com/index/computer-using-agent/. Type: Company research post. | E2 | ||
| 65 | . Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. Type: Company blog. | E3 | ||
| 66 | . SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. Type: Company report. | E2 | ||
| 67 | . Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. Type: Product launch. | E2 |
11. Uncertainty log
This dossier opens no sources. Every sentence is a synthesis from dossiers 01–11 and the 2026-08-28 verification logs. Grades are copied; mixed claims take the lower grade. Not used as warrant: journalism, aggregators, search snippets, unopened full texts, causal acceleration sentences that stand in no opened source.
- Established (layer 1): compute laws (Kaplan/Chinchilla) plus inference energy (Luccioni) plus US DC break (LBNL); agent evals with human gap and scaffold dependence (reopened in 03); HITL task RCTs (Noy, Peng, Brynjolfsson) against Census adoption; AF2/Terwilliger accelerate but do not replace; HELM/SWE as instrument; Pan/Gao classical GPU counters to NISQ advantage.
- Claimed (layer 2): IEA DC path and agent-query energy; HBM/capex; Devin/CUA/computer use with their own caveats; Figure/Atlas demos; Eloundou/Frey exposure; Gidney resource estimate.
- Constrained (layer 3): Jevons net; memory wall; autonomy = HITL; SWE-Verified contamination; exposure ≠ displacement; education see D11 (Bastani opened, PISA unopened); prediction ≠ design ≠ clinic; humanoids not in the IFR stock; QC ≠ sensing; UPS ≠ grid-shift.
Not opened (not a warrant)
- No new source opened in this synthesis run (note 2026-08-31: Epoch report in D01, Masanet in D02, ESMFold bioRxiv in D06 opened — figures there, not copied here; Bastani/ProteinMPNN/Palmer remain in the sister dossiers). Unwarranted couplings: AlphaMissense as LLM-protein (RoseTTAFold opened in D06); RFdiffusion/AF to clinic or syn3.0; quantum ML / ML-for-QEC; quantum sensing as an enabler of FT-QC; ML-designed battery cells; π0/RT-X/Optimus foundation-model embodiment; in-vivo gene therapy; IEA methods annex for the exact order of magnitude/year; Epoch chips-topic-overview; live leaderboards 2026; WebVoyager; agent system cards as PDF.
- Dossier 10 log: SWE-Verified and SWE-Pro are E0 in that run (timeout / sister log). Figures for them appear in this entry only because 01 and 03 reopened them — not because 10 carries them.
- Visual understanding 2026 (GPT-4 technical report unread in 01) stays E0 and is not used here as a multimodal coupling to robotics or computer use.
- IEA 20–25 GW batteries in DCs by 2030 sits in the energy-log table, not as a claim in energie-chips.mjs; not used as a figure here. Warrant for DC batteries remains UPS 45 GW 2025 (GER 2026) as a short-time bridge.