Dossier 03
Agents
v1.4-draft · unchecked
- E1 confirmed
- E2 single-source
- E3 claimed
- E4 disputed
- E0 unknown
1. As of
- Date
- Version
- v1.4-draft
- Author / model
- Atlas generator
- Reviewer
- unchecked (Josef)
2. In one sentence
Execution evals for agents exist (GAIA, τ-bench, OSWorld, WebArena, AgentBench, InterCode); Reflexion improves multi-trial loops without weight updates; Tree of Thoughts searches over intermediate “thoughts” (Game of 24 ToT b=5 74% vs CoT 4%); in the opened paper snapshots models often sit far below humans, consistency (pass^k) collapses, interaction helps but plateaus, and product autonomy is by design not unsupervised — confirmation and watch mode.
Established now · E1 / E2
3. What works today
Reason+Act is a method, not a solved autonomy problem
-
E1
ReAct interleaved thought/action/observation. ALFWorld (PaLM-540B, best-of-6): ReAct 71% vs. Act 45%. WebShop: ReAct success rate 40.0 vs. human 59.6. Few-shot; loops documented; limited action spaces. E1 for the 2022/23 paper measurement; not E1 for “agents are autonomous”.
Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/html/2210.03629. As of 2022-10. Checked 2026-08-28. Type: Paper (ICLR 2023). -
E1
Toolformer (GPT-J 6.7B) learns API calls self-supervised. Explicit limits in paper §7: cannot chain tools; no interactive search; at most one API call per input. Does not beat GPT-3 on all QA sets. This is not an autonomous multi-step agent.
Schick et al. Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/html/2302.04761. As of 2023-02-09. Checked 2026-08-28. Type: Paper.
Classical planning (PlanBench)
-
E1
Autonomous mode, domain in the prompt, VAL validation, T=0. GPT-4 natural-language one-shot Blocksworld 206/600 (34.3%), zero-shot 210/600 (34.6%), CoT 214/600 (35.6%). Logistics one-shot 28/200 (14%). Mystery Blocksworld (deceptive names) one-shot 26/600 (4.3%), zero-shot 1/600. Abstract: best model GPT-4 average success ~12% across domains. Humans on Blocksworld (n=50, IRB): 39/50 (78%) valid, 35/39 of those optimal. Fine-tune GPT-3 on 1000 BW instances: 122/600 (~20%). Obfuscation destroys LLM performance; classical planners are invariant — a hint at pattern matching, not domain independence.
Valmeekam et al. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. As of 2023-05 / NeurIPS 2023 PlanBench. Checked 2026-08-29. Type: Paper. -
E2
LLM-Modulo: GPT-4 seed plans reduce LPG search steps (BW 15.8 empty → 8.9; Logistics 77.5 → 51.3); Mystery no gain. VAL backprompting (max 15 rounds, 50 failed instances): GPT-4 BW 41/50 (82%, avg 3.68 rounds), Logistics 35/50 (70%), Mystery 5/50 (10%). Human+LLM suggestion: no significant accuracy/time/load difference; 3/48 accepted wrong LLM plans. Autonomy ≠ heuristic.
Valmeekam et al. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. As of 2023-05 / NeurIPS 2023 PlanBench. Checked 2026-08-29. Type: Paper.
Execution evals: humans ≫ models in the paper snapshots
-
E1
GAIA: 466 questions. Humans 92%. GPT-4 + plugins (oracle: plugins chosen manually): 30.3% L1 / 9.7% L2 / 0% L3. Abstract “15% GPT-4 plugins” vs. table oracle — both in the paper; the oracle caveat belongs to the claim. Live leaderboard 2026 not opened.
Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper. -
E1
τ-bench function-calling pass^1: gpt-4o 61.2% retail / 35.2% airline / avg 48.2. pass^8 retail gpt-4o <25%. Policy ablation airline gpt-4o 33.2 → 10.8. User sim = LM; reward = DB end state, not sufficient for policy fidelity.
Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.τ-bench, Table 2, function calling, ≥3 trials, as of the 2024 paper. Model retail pass^1 airline pass^1 avg gpt-4o 61.2 35.2 48.2 gpt-4-turbo 57.7 32.4 45.1 claude-3-opus 44.2 34.7 39.5 gpt-3.5-turbo 20.0 10.8 15.4 llama-3-70B (text-ReAct) 14.8 14.4 14.6 -
E1
WebArena: 812 tasks. Best GPT-4 agent 14.41% vs. human 78.24%. GPT-4 flagged 54.9% of feasible tasks as impossible with UA hint.
Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper. -
E1
OSWorld: humans 72.36%. Best in the paper: GPT-4 + accessibility tree 12.24%; screenshot-only GPT-4V 5.26%. Max 15 steps. Later vendor figures (14.9 / 38.1 / 61.4) are other models and step budgets — layer 2.
Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.
Coding agents: scaffold moves the score
-
E1
SWE-bench: 2,294 issues, 12 Python repos. BM25 Claude 2 1.96% resolved; oracle Claude 2 4.80%. Non-interactive RAG, not an agent.
Jimenez et al. SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. As of 2023-10-10. Checked 2026-08-28. Type: Paper. -
E1
SWE-agent (ACI): GPT-4 Turbo full 12.47%, Lite 18.00%. Shell-only Lite 11.00% vs. ACI 18.00%. 51.7% of trajectories have ≥1 failed edit. Scaffold-dependent.
Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper. -
E2
SWE-Bench Pro: gold mean 107.4 LOC / 4.1 files. Public: Claude Sonnet 4.5 43.6%, GPT-5 high 41.8%. Commercial: Claude Opus 4.1 17.8%, GPT-5 high 15.7%. Ablation: GPT-5 high 25.9% with augmentations vs. 8.40% problem-statement-only. One lab, one scaffold.
Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI). -
E1
pass^k falls steeply; spec text and ACI move scores more than “the model alone”. Autonomy scores are scaffold scores.
Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).
Multi-environment agent evals (AgentBench / InterCode)
-
E1
AgentBench (Liu et al., arXiv:2308.03688 / ICLR 2024): 8 environments, 29 LLMs. gpt-4 (0613) overall AgentBench OA 4.01; House-Holding SR 78.0; OS 42.4; DB 32.0; KG 58.8; DCG 74.5; LTP 16.6; WebShop 61.1; Web Browsing 29.0. API-based avg OA 2.32 vs OSS avg 0.51; best OSS ≤70B in scope: CodeLlama-34B-Instruct OA 0.96. Authors: even strongest gpt-4 “not qualified as a practically usable agent”; main obstacles long-term reasoning, decision-making, instruction following; predominant failure Task Limit Exceeded (TLE). CoT-only primitive evaluation (T=0); not multi-trial Reflexion/ToT.
Liu et al. AgentBench. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. As of 2023-08 / ICLR 2024. Checked 2026-09-07. Type: Paper (arXiv; ICLR). -
E1
InterCode (Yang et al., arXiv:2306.14898): interactive coding as RL env (Bash/SQL/Python, Docker). gpt-4 InterCode-SQL: Single Turn All 9.1% → Try Again (n=10) 73.7% success rate; InterCode-Bash: Single Turn All 34.0% → Try Again 48.5%. Interaction helps; late-turn plateau — models less capable as context builds.
Yang et al. InterCode. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. As of 2023-06. Checked 2026-09-07. Type: Paper (arXiv).
Reflexion: verbal reinforcement without weight updates
-
E1
Reflexion (Shinn et al., arXiv:2303.11366, body opened): agent verbally reflects on feedback, stores text in episodic memory; no weight updates. Actor / evaluator / self-reflection.
Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS). -
E1
Paper gains vs strong baselines: AlfWorld +22% absolute (12 iterative steps, 134 envs); HotPotQA +20%; HumanEval Python up to +11%.
Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS). -
E1
Table 1 Pass@1: HumanEval (PY) Reflexion 91.0 vs GPT-4 80.1; HumanEval (RS) 68.0 vs 60.0; Leetcode Hard (PY) 15.0 vs 7.5. MBPP (PY): Reflexion 77.1 vs GPT-4 80.1 — not uniformly better.
Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).
Tree of Thoughts: search over intermediate thoughts
-
E1
Tree of Thoughts (Yao et al., arXiv:2305.10601, body opened): search over intermediate “thoughts”; BFS/DFS; LM proposes and evaluates. Experiments with GPT-4 Chat Completion, sampling temperature 0.7; run May 5–16, 2023. Tasks: Game of 24, Creative Writing, Mini Crosswords.
Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS). -
E1
Game of 24 (100 hard games 901–1000, Table 2): GPT-4 CoT 4.0%; IO 7.3%; CoT-SC k=100 9.0%; ToT b=5 74%; ToT b=1 45%; IO+Refine k=10 27%; CoT best-of-100 49%; IO best-of-100 33%.
Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS). -
E1
Creative Writing and Mini Crosswords are also measured in the paper (ToT above IO/CoT under authors’ protocols); primary figures for this entry from Game of 24 Table 2.
Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).
Claimed · E3
4. What is claimed, not shown
Vendor and product pages 2024–2026
-
E3
Devin: self-description “the first AI software engineer”. E3 (marketing).
Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog. -
E2
Devin SWE-bench 13.86% on a random 25% subset (79/570). Comparison against assisted 4.80%. TDD 23%/100 incomparable (test patch given). E2 for 13.86% as a vendor measurement; E4 as a comparison to “software engineer”.
Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report. -
E2
Anthropic computer use: itself “experimental—at times cumbersome and error-prone.” OSWorld: Claude 3.5 Sonnet 14.9% screenshot-only, with more steps 22.0%. Humans “generally 70–75%”. Prompt-injection risk named. E2 for self-reports; E3 for product utility.
Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.Anthropic research. Developing a computer use model. https://www.anthropic.com/research/developing-computer-use. As of 2024. Checked 2026-08-28. Type: Research post. -
E3
Sonnet 4.5: “best coding model in the world” / “best model at using computers”. E3 (superlatives). Vendor eval: SWE-Verified 77.2% with prompt addendum (“use tools >100 times; write tests first”); OSWorld-Verified 61.4%, 100 max steps. “maintaining focus for more than 30 hours” = observation, not a public eval. System card not opened.
Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news. -
E2
OpenAI CUA: OSWorld 38.1% vs. human 72.4%; WebArena 58.1% vs. human 78.2%. “we don’t expect CUA to perform reliably in all scenarios just yet.” Sensitive actions: user confirmation; watch mode. Operator internal mini-trials n=10: tagvenue with hints 8/10 vs. without 3/10. System card PDF not opened.
OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post. -
E2
ChatGPT agent: permission before consequential actions; watch mode; refuse bank transfers. “still in its early stages”. HLE pass@1 41.6; SpreadsheetBench 45.54% vs. human 71.33%. WebArena percentage in the extracted text not isolated — do not cite as a figure. Launch post itself: outdated. System card not opened.
OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch. -
E2
OpenAI 2026-02-23: SWE-bench Verified no longer measures frontier coding. SOTA 74.9% → 80.9%. Audit of 138 unsolved: 59.4% material test issues. Contamination: GPT-5.2, Claude Opus 4.5, Gemini 3 Flash Preview reproduced gold patches. Sonnet 4.5’s 77.2% sits before this withdrawal.
OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.
Constrained · Limit
5. Bottleneck and limit
Where autonomy becomes illusion
-
E1
Autonomous plan generation on IPC-like domains is unsolved for GPT-4 as of 2023 (~12% avg, Mystery <5%). Product “agents plan” without VAL/LPG remains E3.
Valmeekam et al. On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. As of 2023-05 / NeurIPS 2023 PlanBench. Checked 2026-08-29. Type: Paper. -
E1
The human–model gap on execution evals remains the most robust pattern 2023–2025. GAIA L3 paper: 0%. WebArena paper 14.41% vs. 78%. OSWorld paper 12.24% vs. 72%. The gap closes on some vendor evals (OSWorld 2025), not on the GAIA L3 paper and not on SWE-Pro commercial (17.8%).
Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper. -
E2
“Autonomous” in products is by design not unsupervised: confirmation, watch mode, decline banking. The product model is human-in-the-loop.
Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch. -
E1
Scaffold inflation: GAIA plugins = manually chosen oracle plugins. SWE-agent ACI vs. shell 18 vs. 11. SWE-Pro 25.9 vs. 8.40. Sonnet 4.5 prompt addendum. Operator tagvenue 8/10 vs. 3/10. A dossier sentence without these constraints would be false.
Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post. -
E2
Contamination and broken tests make high coding scores unreadable. SWE-Pro commercial ≪ public (17.8 vs. 43.6). Contamination resistance in the Pro paper: E3 for the resistance, E2 for the split.
OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI). -
E1
AgentBench OA is a weighted snapshot (Weight-1 from tested LLMs) — not a 2026 live leaderboard and not a universal capability score. House-Holding 78% is ALFWorld SR under the AgentBench protocol, not factory autonomy.
Liu et al. AgentBench. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. As of 2023-08 / ICLR 2024. Checked 2026-09-07. Type: Paper (arXiv; ICLR). -
E1
InterCode measures interactive coding with execution feedback in Docker — not evidence that interactive agents are production software engineers.
Yang et al. InterCode. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. As of 2023-06. Checked 2026-09-07. Type: Paper (arXiv). -
E1
Reflexion is multi-trial verbal reinforcement (memory), not single-shot and not a weight update. HumanEval 91% under Reflexion protocol ≠ “agents solved”; MBPP PY 77.1 < GPT-4 80.1.
Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS). -
E1
ToT is deliberate search over intermediate thoughts — not autonomy. GPT-4 Chat Completion T=0.7, May 2023; Game of 24 ≠ general agents. Creative Writing / Mini Crosswords are further paper tasks, not production agents.
Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS).
6. Actors and incentives
Who measures, who sells
-
E1
GAIA, τ-bench, OSWorld, WebArena, AgentBench, InterCode, Reflexion, ToT are paper evals (AgentBench/InterCode without a human baseline in the same sense). Live boards 2026 not opened.
Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.Yao, Shinn et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. As of 2024-06-17. Checked 2026-08-28. Type: Paper.Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.Zhou et al. WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. As of 2023-07. Checked 2026-08-28. Type: Paper.Liu et al. AgentBench. AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. As of 2023-08 / ICLR 2024. Checked 2026-09-07. Type: Paper (arXiv; ICLR).Yang et al. InterCode. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. As of 2023-06. Checked 2026-09-07. Type: Paper (arXiv).Shinn et al. Reflexion. Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. As of 2023-03 / NeurIPS 2023. Checked 2026-09-09. Type: Paper (arXiv; NeurIPS).Yao et al. ToT. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. As of 2023-05 / NeurIPS 2023. Checked 2026-09-10. Type: Paper (arXiv; NeurIPS). -
E2
Cognition, Anthropic, OpenAI sell autonomy on product pages that also carry caveats (error-prone, confirmation, early stages).
Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report.Anthropic computer use. Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. As of 2024-10-22. Checked 2026-08-28. Type: Company news.Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.OpenAI ChatGPT agent. Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. As of 2025-07-17. Checked 2026-08-28. Type: Product launch. -
E2
OpenAI withdraws SWE-Verified in 2026 as a frontier measure; Scale AI publishes SWE-Pro with a public/commercial split.
OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report.Deng, Da et al. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. As of 2025-09. Checked 2026-08-28. Type: Paper (Scale AI).
7. State of the dispute
Marketing vs. measured gap
-
E2
“First AI software engineer” (Devin) stands against 13.86% on a 25% subset, in the same ballpark as SWE-agent 12.47% full.
Cognition Devin. Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. As of 2024. Checked 2026-08-28. Type: Company blog.Cognition SWE-bench. SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. As of 2024-03-15. Checked 2026-08-28. Type: Company report.Yang et al. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. As of 2024-05-06. Checked 2026-08-28. Type: Paper. -
E2
Sonnet 4.5 77.2% SWE-Verified before the OpenAI audit vs. Verified withdrawn as a frontier measure. Without the scaffold footnote, E4.
Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news.OpenAI. Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. As of 2026-02-23. Checked 2026-08-28. Type: Company report. -
E2
OSWorld paper 12.24% (15 steps) vs. CUA 38.1% vs. Sonnet 4.5 OSWorld-Verified 61.4% (100 steps) — not the same protocol.
Xie et al. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. As of 2024-04. Checked 2026-08-28. Type: Paper.OpenAI CUA. Computer-Using Agent. https://openai.com/index/computer-using-agent/. As of 2025-01-23. Checked 2026-08-28. Type: Company research post.Anthropic Sonnet 4.5. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. As of 2025-09-29. Checked 2026-08-28. Type: Company news. -
E0
Live leaderboards 2026 (GAIA HF, OSWorld-Verified, WebArena) not opened — no ranking dispute, only the gap.
Mialon et al. GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. As of 2023-11-21. Checked 2026-08-28. Type: Paper.
8. Open questions
- AgentBench / InterCode — opened 2026-09-07 as sources 19–20; Reflexion 2026-09-09 source 21; ToT 2026-09-10 source 22; live boards and later replications not opened.
- Operator system card PDF and ChatGPT agent system card PDF — preparedness, watch mode, prompt-injection evals.
- GAIA Hugging Face leaderboard 2026 — whether scaffolds close the human gap.
- OSWorld-Verified / OSWorld live — Sonnet 4.5 61.4% and CUA 38.1% are not the same protocol.
- WebVoyager original paper — 87% without a definition of “simple”.
- Anthropic Opus 4.5 / Sonnet 4.6 system cards — search snippets (OSWorld 66.3 / 72.5) unused until the PDF is opened.
- τ-bench / τ2-bench independent replication 2025–2026.
- The Agent Company / VisualWebArena.
- Independent Devin eval outside the Cognition subset.
- What is actually shipped as of 2026-08-28 — ChatGPT agent page itself: launch post outdated.
9. Changes
- v1.4-draft2026-09-10: Tree of Thoughts (Yao et al., arXiv:2305.10601, source 22, E1) opened — Game of 24 ToT b=5 74% vs CoT 4%; deliberate search ≠ autonomy.
- v1.3-draft2026-09-09: Reflexion (Shinn et al., arXiv:2303.11366, source 21, E1) opened — verbal reinforcement without weight update; HumanEval 91% multi-trial; MBPP not uniformly better.
- v1.2-draft2026-09-07: AgentBench (Liu et al., source 19, E1) and InterCode (Yang et al., source 20, E1) opened — multi-env agent evals; OA ≠ live boards; interaction helps, plateaus.
- v1.1-draft2026-08-29: Valmeekam et al. (arXiv:2305.15771 HTML, PlanBench) opened; autonomous planning GPT-4 ~12% avg.
- v1.0-draftFirst version from the 2026-08-28 verification log.
10. Sources
| No. | Source | As of | Checked | Grade |
|---|---|---|---|---|
| 1 | . ReAct: Synergizing Reasoning and Acting in Language Models. https://arxiv.org/html/2210.03629. Type: Paper (ICLR 2023). | E1 | ||
| 2 | . Toolformer: Language Models Can Teach Themselves to Use Tools. https://arxiv.org/html/2302.04761. Type: Paper. | E1 | ||
| 3 | . GAIA: a Benchmark for General AI Assistants. https://arxiv.org/html/2311.12983. Type: Paper. | E1 | ||
| 4 | . τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. https://arxiv.org/html/2406.12045. Type: Paper. | E1 | ||
| 5 | . SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. https://arxiv.org/html/2310.06770. Type: Paper. | E1 | ||
| 6 | . SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. https://arxiv.org/html/2405.15793. Type: Paper. | E1 | ||
| 7 | . SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?. https://arxiv.org/html/2509.16941. Type: Paper (Scale AI). | E2 | ||
| 8 | . OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. https://arxiv.org/html/2404.07972. Type: Paper. | E1 | ||
| 9 | . WebArena: A Realistic Web Environment for Building Autonomous Agents. https://arxiv.org/html/2307.13854. Type: Paper. | E1 | ||
| 10 | . Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku. https://www.anthropic.com/news/3-5-models-and-computer-use. Type: Company news. | E2 | ||
| 11 | . Developing a computer use model. https://www.anthropic.com/research/developing-computer-use. Type: Research post. | E2 | ||
| 12 | . Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5. Type: Company news. | E2 | ||
| 13 | . Introducing Devin, the first AI software engineer. https://cognition.com/blog/introducing-devin. Type: Company blog. | E3 | ||
| 14 | . SWE-bench Technical Report. https://cognition.com/blog/swe-bench-technical-report. Type: Company report. | E2 | ||
| 15 | . Computer-Using Agent. https://openai.com/index/computer-using-agent/. Type: Company research post. | E2 | ||
| 16 | . Introducing ChatGPT agent: bridging research and action. https://openai.com/index/introducing-chatgpt-agent. Type: Product launch. | E2 | ||
| 17 | . Why SWE-bench Verified no longer measures frontier coding capabilities. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/. Type: Company report. | E2 | ||
| 18 | . On the Planning Abilities of Large Language Models: A Critical Investigation. https://arxiv.org/html/2305.15771. Type: Paper. | E1 | ||
| 19 | . AgentBench: Evaluating LLMs as Agents. https://arxiv.org/pdf/2308.03688. Type: Paper (arXiv; ICLR). | E1 | ||
| 20 | . InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback. https://arxiv.org/pdf/2306.14898. Type: Paper (arXiv). | E1 | ||
| 21 | . Reflexion: Language Agents with Verbal Reinforcement Learning. https://arxiv.org/pdf/2303.11366. Type: Paper (arXiv; NeurIPS). | E1 | ||
| 22 | . Tree of Thoughts: Deliberate Problem Solving with Large Language Models. https://arxiv.org/pdf/2305.10601. Type: Paper (arXiv; NeurIPS). | E1 |
11. Uncertainty log
Overall uncertainty of this entry, bound to the verification log of 2026-08-28 plus PlanBench 2026-08-29 plus AgentBench/InterCode 2026-09-07 plus Reflexion 2026-09-09 plus ToT 2026-09-10. 22 full-text openings, 2 partial openings, at least 15 searched but not opened. The dossier-1 log is not a warrant — figures only where the primary source was re-opened here.
- Established (layer 1): ReAct, Toolformer, Reflexion and ToT as methods with documented limits; PlanBench/classical planning (GPT-4 ~12% avg autonomous); GAIA, τ-bench, WebArena, OSWorld, SWE-bench/SWE-agent, AgentBench (OA 4.01 / TLE), InterCode (SQL 9.1→73.7%); scaffold dependence.
- Claimed (layer 2): Devin, computer use, Sonnet 4.5, CUA/Operator, ChatGPT agent — vendor figures with their own caveats.
- Constrained (layer 3): autonomous IPC planning without VAL/LPG; AgentBench OA ≠ 2026 live boards; HH 78% ≠ factory; InterCode ≠ production SWE; GUI grounding; prompt injection (measured ASR 2026 E0); watch mode; SWE-Verified contamination; live leaderboards 2026.
Not opened (not a warrant)
- AgentBench / InterCode — opened 2026-09-07 (sources 19–20); Reflexion opened 2026-09-09 (source 21); ToT opened 2026-09-10 (source 22); live boards and later replications unopened.
- OpenAI Operator system card PDF — WebFetch 429.
- ChatGPT agent system card — timeout / shell reject.
- Anthropic Opus 4.5 / Sonnet 4.6 system cards — search snippets (OSWorld 66.3%, SWE-Verified 80.9%) = E0.
- WebVoyager original paper — cited only via the CUA page (87%).
- VisualWebArena; AutoGPT paper; The Agent Company — not opened. Reflexion opened 2026-09-09 (source 21); ToT opened 2026-09-10 (source 22).
- GAIA / OSWorld / WebArena live leaderboards 2026 — not opened.
- ChatGPT agent WebArena % in extracted HTML not isolated — not a figure.