Atlas of the Present Atlas · v0.1

As of 12 September 2026

Dossier 03

Agents

v1.4-draft · unchecked

1. As of

Date
Version
v1.4-draft
Author / model
Atlas generator
Reviewer
unchecked (Josef)

2. In one sentence

Execution evals for agents exist (GAIA, τ-bench, OSWorld, WebArena, AgentBench, InterCode); Reflexion improves multi-trial loops without weight updates; Tree of Thoughts searches over intermediate “thoughts” (Game of 24 ToT b=5 74% vs CoT 4%); in the opened paper snapshots models often sit far below humans, consistency (pass^k) collapses, interaction helps but plateaus, and product autonomy is by design not unsupervised — confirmation and watch mode.

Established now · E1 / E2

3. What works today

Reason+Act is a method, not a solved autonomy problem

  1. E1

    ReAct interleaved thought/action/observation. ALFWorld (PaLM-540B, best-of-6): ReAct 71% vs. Act 45%. WebShop: ReAct success rate 40.0 vs. human 59.6. Few-shot; loops documented; limited action spaces. E1 for the 2022/23 paper measurement; not E1 for “agents are autonomous”.

    Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/html/2210.03629. As of 2022-10. Checked 2026-08-28. Type: Paper (ICLR 2023).
  2. E1

    Toolformer (GPT-J 6.7B) learns API calls self-supervised. Explicit limits in paper §7: cannot chain tools; no interactive search; at most one API call per input. Does not beat GPT-3 on all QA sets. This is not an autonomous multi-step agent.

    Schick et al. Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/html/2302.04761. As of 2023-02-09. Checked 2026-08-28. Type: Paper.

Classical planning (PlanBench)

  1. E1

    Autonomous mode, domain in the prompt, VAL validation, T=0. GPT-4 natural-language one-shot Blocksworld 206/600 (34.3%), zero-shot 210/600 (34.6%), CoT 214/600 (35.6%). Logistics one-shot 28/200 (14%). Mystery Blocksworld (deceptive names) one-shot 26/600 (4.3%), zero-shot 1/600. Abstract: best model GPT-4 average success ~12% across domains. Humans on Blocksworld (n=50, IRB): 39/50 (78%) valid, 35/39 of those optimal. Fine-tune GPT-3 on 1000 BW instances: 122/600 (~20%). Obfuscation destroys LLM performance; classical planners are invariant — a hint at pattern matching, not domain independence.

    Valmeekam et al. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. As of 2023-05 / NeurIPS 2023 PlanBench. Checked 2026-08-29. Type: Paper.
  2. E2

    LLM-Modulo: GPT-4 seed plans reduce LPG search steps (BW 15.8 empty → 8.9; Logistics 77.5 → 51.3); Mystery no gain. VAL backprompting (max 15 rounds, 50 failed instances): GPT-4 BW 41/50 (82%, avg 3.68 rounds), Logistics 35/50 (70%), Mystery 5/50 (10%). Human+LLM suggestion: no significant accuracy/time/load difference; 3/48 accepted wrong LLM plans. Autonomy ≠ heuristic.

    Valmeekam et al. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. As of 2023-05 / NeurIPS 2023 PlanBench. Checked 2026-08-29. Type: Paper.

Execution evals: humans ≫ models in the paper snapshots

  1. E1

    GAIA: 466 questions. Humans 92%. GPT-4 + plugins (oracle: plugins chosen manually): 30.3% L1 / 9.7% L2 / 0% L3. Abstract “15% GPT-4 plugins” vs. table oracle — both in the paper; the oracle caveat belongs to the claim. Live leaderboard 2026 not opened.

    Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.
  2. E1

    τ-bench function-calling pass^1: gpt-4o 61.2% retail / 35.2% airline / avg 48.2. pass^8 retail gpt-4o <25%. Policy ablation airline gpt-4o 33.2 → 10.8. User sim = LM; reward = DB end state, not sufficient for policy fidelity.

    τ-bench, Table 2, function calling, ≥3 trials, as of the 2024 paper.
    Modelretail pass^1airline pass^1avg
    gpt-4o61.235.248.2
    gpt-4-turbo57.732.445.1
    claude-3-opus44.234.739.5
    gpt-3.5-turbo20.010.815.4
    llama-3-70B (text-ReAct)14.814.414.6
    Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.
  3. E1

    WebArena: 812 tasks. Best GPT-4 agent 14.41% vs. human 78.24%. GPT-4 flagged 54.9% of feasible tasks as impossible with UA hint.

    Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper.
  4. E1

    OSWorld: humans 72.36%. Best in the paper: GPT-4 + accessibility tree 12.24%; screenshot-only GPT-4V 5.26%. Max 15 steps. Later vendor figures (14.9 / 38.1 / 61.4) are other models and step budgets — layer 2.

    Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.

Coding agents: scaffold moves the score

  1. E1

    SWE-bench: 2,294 issues, 12 Python repos. BM25 Claude 2 1.96% resolved; oracle Claude 2 4.80%. Non-interactive RAG, not an agent.

    Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10-10. Checked 2026-08-28. Type: Paper.
  2. E1

    SWE-agent (ACI): GPT-4 Turbo full 12.47%, Lite 18.00%. Shell-only Lite 11.00% vs. ACI 18.00%. 51.7% of trajectories have ≥1 failed edit. Scaffold-dependent.

    Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.
  3. E2

    SWE-Bench Pro: gold mean 107.4 LOC / 4.1 files. Public: Claude Sonnet 4.5 43.6%, GPT-5 high 41.8%. Commercial: Claude Opus 4.1 17.8%, GPT-5 high 15.7%. Ablation: GPT-5 high 25.9% with augmentations vs. 8.40% problem-statement-only. One lab, one scaffold.

    Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).
  4. E1

    pass^k falls steeply; spec text and ACI move scores more than “the model alone”. Autonomy scores are scaffold scores.

    Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).

Multi-environment agent evals (AgentBench / InterCode)

  1. E1

    AgentBench (Liu et al., arXiv:2308.03688 / ICLR 2024): 8 environments, 29 LLMs. gpt-4 (0613) overall AgentBench OA 4.01; House-Holding SR 78.0; OS 42.4; DB 32.0; KG 58.8; DCG 74.5; LTP 16.6; WebShop 61.1; Web Browsing 29.0. API-based avg OA 2.32 vs OSS avg 0.51; best OSS ≤70B in scope: CodeLlama-34B-Instruct OA 0.96. Authors: even strongest gpt-4 “not qualified as a practically usable agent”; main obstacles long-term reasoning, decision-making, instruction following; predominant failure Task Limit Exceeded (TLE). CoT-only primitive evaluation (T=0); not multi-trial Reflexion/ToT.

    Liu et al. AgentBench. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. As of 2023-08 / ICLR 2024. Checked 2026-09-07. Type: Paper (arXiv; ICLR).
  2. E1

    InterCode (Yang et al., arXiv:2306.14898): interactive coding as RL env (Bash/SQL/Python, Docker). gpt-4 InterCode-SQL: Single Turn All 9.1% → Try Again (n=10) 73.7% success rate; InterCode-Bash: Single Turn All 34.0% → Try Again 48.5%. Interaction helps; late-turn plateau — models less capable as context builds.

    Yang et al. InterCode. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. As of 2023-06. Checked 2026-09-07. Type: Paper (arXiv).

Reflexion: verbal reinforcement without weight updates

  1. E1

    Reflexion (Shinn et al., arXiv:2303.11366, body opened): agent verbally reflects on feedback, stores text in episodic memory; no weight updates. Actor / evaluator / self-reflection.

    Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).
  2. E1

    Paper gains vs strong baselines: AlfWorld +22% absolute (12 iterative steps, 134 envs); HotPotQA +20%; HumanEval Python up to +11%.

    Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).
  3. E1

    Table 1 Pass@1: HumanEval (PY) Reflexion 91.0 vs GPT-4 80.1; HumanEval (RS) 68.0 vs 60.0; Leetcode Hard (PY) 15.0 vs 7.5. MBPP (PY): Reflexion 77.1 vs GPT-4 80.1 — not uniformly better.

    Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).

Tree of Thoughts: search over intermediate thoughts

  1. E1

    Tree of Thoughts (Yao et al., arXiv:2305.10601, body opened): search over intermediate “thoughts”; BFS/DFS; LM proposes and evaluates. Experiments with GPT-4 Chat Completion, sampling temperature 0.7; run May 5–16, 2023. Tasks: Game of 24, Creative Writing, Mini Crosswords.

    Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).
  2. E1

    Game of 24 (100 hard games 901–1000, Table 2): GPT-4 CoT 4.0%; IO 7.3%; CoT-SC k=100 9.0%; ToT b=5 74%; ToT b=1 45%; IO+Refine k=10 27%; CoT best-of-100 49%; IO best-of-100 33%.

    Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).
  3. E1

    Creative Writing and Mini Crosswords are also measured in the paper (ToT above IO/CoT under authors’ protocols); primary figures for this entry from Game of 24 Table 2.

    Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).

Claimed · E3

4. What is claimed, not shown

Vendor and product pages 2024–2026

  1. E3

    Devin: self-description “the first AI software engineer”. E3 (marketing).

    Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.
  2. E2

    Devin SWE-bench 13.86% on a random 25% subset (79/570). Comparison against assisted 4.80%. TDD 23%/100 incomparable (test patch given). E2 for 13.86% as a vendor measurement; E4 as a comparison to “software engineer”.

    Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report.
  3. E2

    Anthropic computer use: itself “experimental—at times cumbersome and error-prone.” OSWorld: Claude 3.5 Sonnet 14.9% screenshot-only, with more steps 22.0%. Humans “generally 70–75%”. Prompt-injection risk named. E2 for self-reports; E3 for product utility.

    Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.Anthropic research. Developing a computer use model. https://www.anthropic.com/research/developing-computer-use. As of 2024. Checked 2026-08-28. Type: Research post.
  4. E3

    Sonnet 4.5: “best coding model in the world” / “best model at using computers”. E3 (superlatives). Vendor eval: SWE-Verified 77.2% with prompt addendum (“use tools >100 times; write tests first”); OSWorld-Verified 61.4%, 100 max steps. “maintaining focus for more than 30 hours” = observation, not a public eval. System card not opened.

    Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.
  5. E2

    OpenAI CUA: OSWorld 38.1% vs. human 72.4%; WebArena 58.1% vs. human 78.2%. “we don’t expect CUA to perform reliably in all scenarios just yet.” Sensitive actions: user confirmation; watch mode. Operator internal mini-trials n=10: tagvenue with hints 8/10 vs. without 3/10. System card PDF not opened.

    OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.
  6. E2

    ChatGPT agent: permission before consequential actions; watch mode; refuse bank transfers. “still in its early stages”. HLE pass@1 41.6; SpreadsheetBench 45.54% vs. human 71.33%. WebArena percentage in the extracted text not isolated — do not cite as a figure. Launch post itself: outdated. System card not opened.

    OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch.
  7. E2

    OpenAI 2026-02-23: SWE-bench Verified no longer measures frontier coding. SOTA 74.9% → 80.9%. Audit of 138 unsolved: 59.4% material test issues. Contamination: GPT-5.2, Claude Opus 4.5, Gemini 3 Flash Preview reproduced gold patches. Sonnet 4.5’s 77.2% sits before this withdrawal.

    OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.

Constrained · Limit

5. Bottleneck and limit

Where autonomy becomes illusion

  1. E1

    Autonomous plan generation on IPC-like domains is unsolved for GPT-4 as of 2023 (~12% avg, Mystery <5%). Product “agents plan” without VAL/LPG remains E3.

    Valmeekam et al. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. As of 2023-05 / NeurIPS 2023 PlanBench. Checked 2026-08-29. Type: Paper.
  2. E1

    The human–model gap on execution evals remains the most robust pattern 2023–2025. GAIA L3 paper: 0%. WebArena paper 14.41% vs. 78%. OSWorld paper 12.24% vs. 72%. The gap closes on some vendor evals (OSWorld 2025), not on the GAIA L3 paper and not on SWE-Pro commercial (17.8%).

    Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper.
  3. E2

    “Autonomous” in products is by design not unsupervised: confirmation, watch mode, decline banking. The product model is human-in-the-loop.

    Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch.
  4. E1

    Scaffold inflation: GAIA plugins = manually chosen oracle plugins. SWE-agent ACI vs. shell 18 vs. 11. SWE-Pro 25.9 vs. 8.40. Sonnet 4.5 prompt addendum. Operator tagvenue 8/10 vs. 3/10. A dossier sentence without these constraints would be false.

    Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.
  5. E2

    Contamination and broken tests make high coding scores unreadable. SWE-Pro commercial ≪ public (17.8 vs. 43.6). Contamination resistance in the Pro paper: E3 for the resistance, E2 for the split.

    OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).
  6. E1

    AgentBench OA is a weighted snapshot (Weight-1 from tested LLMs) — not a 2026 live leaderboard and not a universal capability score. House-Holding 78% is ALFWorld SR under the AgentBench protocol, not factory autonomy.

    Liu et al. AgentBench. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. As of 2023-08 / ICLR 2024. Checked 2026-09-07. Type: Paper (arXiv; ICLR).
  7. E1

    InterCode measures interactive coding with execution feedback in Docker — not evidence that interactive agents are production software engineers.

    Yang et al. InterCode. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. As of 2023-06. Checked 2026-09-07. Type: Paper (arXiv).
  8. E1

    Reflexion is multi-trial verbal reinforcement (memory), not single-shot and not a weight update. HumanEval 91% under Reflexion protocol ≠ “agents solved”; MBPP PY 77.1 < GPT-4 80.1.

    Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).
  9. E1

    ToT is deliberate search over intermediate thoughts — not autonomy. GPT-4 Chat Completion T=0.7, May 2023; Game of 24 ≠ general agents. Creative Writing / Mini Crosswords are further paper tasks, not production agents.

    Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).

6. Actors and incentives

Who measures, who sells

  1. E1

    GAIA, τ-bench, OSWorld, WebArena, AgentBench, InterCode, Reflexion, ToT are paper evals (AgentBench/InterCode without a human baseline in the same sense). Live boards 2026 not opened.

    Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper.Liu et al. AgentBench. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. As of 2023-08 / ICLR 2024. Checked 2026-09-07. Type: Paper (arXiv; ICLR).Yang et al. InterCode. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. As of 2023-06. Checked 2026-09-07. Type: Paper (arXiv).Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).
  2. E2

    Cognition, Anthropic, OpenAI sell autonomy on product pages that also carry caveats (error-prone, confirmation, early stages).

    Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report.Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch.
  3. E2

    OpenAI withdraws SWE-Verified in 2026 as a frontier measure; Scale AI publishes SWE-Pro with a public/commercial split.

    OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).

7. State of the dispute

Marketing vs. measured gap

  1. E2

    “First AI software engineer” (Devin) stands against 13.86% on a 25% subset, in the same ballpark as SWE-agent 12.47% full.

    Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.
  2. E2

    Sonnet 4.5 77.2% SWE-Verified before the OpenAI audit vs. Verified withdrawn as a frontier measure. Without the scaffold footnote, E4.

    Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.
  3. E2

    OSWorld paper 12.24% (15 steps) vs. CUA 38.1% vs. Sonnet 4.5 OSWorld-Verified 61.4% (100 steps) — not the same protocol.

    Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.
  4. E0

    Live leaderboards 2026 (GAIA HF, OSWorld-Verified, WebArena) not opened — no ranking dispute, only the gap.

    Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.

8. Open questions

  1. AgentBench / InterCode — opened 2026-09-07 as sources 19–20; Reflexion 2026-09-09 source 21; ToT 2026-09-10 source 22; live boards and later replications not opened.
  2. Operator system card PDF and ChatGPT agent system card PDF — preparedness, watch mode, prompt-injection evals.
  3. GAIA Hugging Face leaderboard 2026 — whether scaffolds close the human gap.
  4. OSWorld-Verified / OSWorld live — Sonnet 4.5 61.4% and CUA 38.1% are not the same protocol.
  5. WebVoyager original paper — 87% without a definition of “simple”.
  6. Anthropic Opus 4.5 / Sonnet 4.6 system cards — search snippets (OSWorld 66.3 / 72.5) unused until the PDF is opened.
  7. τ-bench / τ2-bench independent replication 2025–2026.
  8. The Agent Company / VisualWebArena.
  9. Independent Devin eval outside the Cognition subset.
  10. What is actually shipped as of 2026-08-28 — ChatGPT agent page itself: launch post outdated.

9. Changes

  • v1.4-draft2026-09-10: Tree of Thoughts (Yao et al., arXiv:2305.10601, source 22, E1) opened — Game of 24 ToT b=5 74% vs CoT 4%; deliberate search ≠ autonomy.
  • v1.3-draft2026-09-09: Reflexion (Shinn et al., arXiv:2303.11366, source 21, E1) opened — verbal reinforcement without weight update; HumanEval 91% multi-trial; MBPP not uniformly better.
  • v1.2-draft2026-09-07: AgentBench (Liu et al., source 19, E1) and InterCode (Yang et al., source 20, E1) opened — multi-env agent evals; OA ≠ live boards; interaction helps, plateaus.
  • v1.1-draft2026-08-29: Valmeekam et al. (arXiv:2305.15771 HTML, PlanBench) opened; autonomous planning GPT-4 ~12% avg.
  • v1.0-draftFirst version from the 2026-08-28 verification log.

10. Sources

No. Source As of Checked Grade
1 Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y.. ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/html/2210.03629. Type: Paper (ICLR 2023). E1
2 Schick, T.; Dwivedi-Yu, J.; Dessì, R.; Raileanu, R.; Lomeli, M.; Zettlemoyer, L.; Cancedda, N.; Scialom, T.. Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/html/2302.04761. Type: Paper. E1
3 Mialon, G.; Fourrier, C.; Swift, C.; Wolf, T.; LeCun, Y.; Scialom, T.. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. Type: Paper. E1
4 Yao, S.; Shinn, N.; Razavi, P.; Narasimhan, K.. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. Type: Paper. E1
5 Jimenez, C. E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K.. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. Type: Paper. E1
6 Yang, J.; Jimenez, C. E.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.; Press, O.. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. Type: Paper. E1
7 Deng, X.; Da, J.; et al. (Scale AI). SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. Type: Paper (Scale AI). E2
8 Xie, T.; Zhang, D.; Chen, J.; Li, X.; Zhao, S.; Cao, R.; et al.. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. Type: Paper. E1
9 Zhou, S.; Xu, F. F.; Zhu, H.; Zhou, X.; Lo, R.; Sridhar, A.; Cheng, X.; Ou, T.; Bisk, Y.; Fried, D.; Alon, U.; Neubig, G.. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. Type: Paper. E1
10 Anthropic. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. Type: Company news. E2
11 Anthropic. Developing a computer use model. https://www.anthropic.com/research/developing-computer-use. Type: Research post. E2
12 Anthropic. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. Type: Company news. E2
13 Cognition. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. Type: Company blog. E3
14 Cognition. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. Type: Company report. E2
15 OpenAI. Computer-Using Agent. https://openai.com/index/computer-using-agent/. Type: Company research post. E2
16 OpenAI. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. Type: Product launch. E2
17 OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. Type: Company report. E2
18 Valmeekam, K.; Marquez, M.; Sreedharan, S.; Kambhampati, S.. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. Type: Paper. E1
19 Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; et al.. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. Type: Paper (arXiv; ICLR). E1
20 Yang, J.; Prabhakar, A.; Narasimhan, K.; Yao, S.. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. Type: Paper (arXiv). E1
21 Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; Yao, S.. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. Type: Paper (arXiv; NeurIPS). E1
22 Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T. L.; Cao, Y.; Narasimhan, K.. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. Type: Paper (arXiv; NeurIPS). E1

11. Uncertainty log

Overall uncertainty of this entry, bound to the verification log of 2026-08-28 plus PlanBench 2026-08-29 plus AgentBench/InterCode 2026-09-07 plus Reflexion 2026-09-09 plus ToT 2026-09-10. 22 full-text openings, 2 partial openings, at least 15 searched but not opened. The dossier-1 log is not a warrant — figures only where the primary source was re-opened here.

  • Established (layer 1): ReAct, Toolformer, Reflexion and ToT as methods with documented limits; PlanBench/classical planning (GPT-4 ~12% avg autonomous); GAIA, τ-bench, WebArena, OSWorld, SWE-bench/SWE-agent, AgentBench (OA 4.01 / TLE), InterCode (SQL 9.1→73.7%); scaffold dependence.
  • Claimed (layer 2): Devin, computer use, Sonnet 4.5, CUA/Operator, ChatGPT agent — vendor figures with their own caveats.
  • Constrained (layer 3): autonomous IPC planning without VAL/LPG; AgentBench OA ≠ 2026 live boards; HH 78% ≠ factory; InterCode ≠ production SWE; GUI grounding; prompt injection (measured ASR 2026 E0); watch mode; SWE-Verified contamination; live leaderboards 2026.

Not opened (not a warrant)

  • AgentBench / InterCode — opened 2026-09-07 (sources 19–20); Reflexion opened 2026-09-09 (source 21); ToT opened 2026-09-10 (source 22); live boards and later replications unopened.
  • OpenAI Operator system card PDF — WebFetch 429.
  • ChatGPT agent system card — timeout / shell reject.
  • Anthropic Opus 4.5 / Sonnet 4.6 system cards — search snippets (OSWorld 66.3%, SWE-Verified 80.9%) = E0.
  • WebVoyager original paper — cited only via the CUA page (87%).
  • VisualWebArena; AutoGPT paper; The Agent Company — not opened. Reflexion opened 2026-09-09 (source 21); ToT opened 2026-09-10 (source 22).
  • GAIA / OSWorld / WebArena live leaderboards 2026 — not opened.
  • ChatGPT agent WebArena % in extracted HTML not isolated — not a figure.