Atlas of the Present Atlas · v0.1

As of 14 September 2026

Not the record

Continuations

The atlas states what is evidenced today. Here stands only what can follow logically — not what will happen.

Each card: Today · What follows · What this means for you. Technical fields stay collapsed.

F1 F2 F3 F4 F5 F6 F8
  1. F1 constructed (E0)

    The work setup makes the difference

    Today

    When you measure coding agents, you often measure the setup around the model, not the model alone. Many runs fail mid-way on code changes.

    What follows

    The result comes from the setup, not from the bare model.

    What this means for you

    • The machine can take the first step. Starting small code changes is easier with the helper setup.
    • Checking and deciding stay with you. You check failed mid-steps, not only the final result.
    • This is what to practice. Set the result with the setup beside the result without it.

    Not claimed The setup makes code changes reliable.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. SWE-agent (ACI): GPT-4 Turbo Lite 18.00 % vs. Shell-only Lite 11.00 %; 51.7 % der Trajectories haben ≥1 failed edit; Scaffold-abhängig. Agents E1
    2. Coding-Agenten: der Scaffold bewegt den Score stärker als „das Modell allein“; Autonomie-Scores sind Scaffold-Scores. Linkages E1
    Which rule (Operation)
    scaffold-nicht-modell
    What follows (Consequence)
    Dieselbe Ausführungsmessung bleibt an ACI/Spec-Text und Scaffold gebunden, nicht an das nackte Modell.
    How far it holds (Limit)
    SWE-agent ACI vs. Shell-only Lite 18.00 % vs. 11.00 %; Scaffold-Inflation: Spec-Text und ACI bewegen Scores stärker als das Modell allein. Autonomie-Scores sind Scaffold-Scores.
    When it tips (Tipping condition)
    Auf demselben SWE-agent-Lite- oder SWE-Pro-Ablation-Protokoll liegt der Score ohne ACI/Augmentation nicht mehr unter dem Scaffold-Lauf.
  2. F2 constructed (E0)

    Oversight stays in the product

    Today

    Products still ask you to confirm. Measured work still sits far behind a human.

    What follows

    The gap remains. The product does not leave the agent alone.

    What this means for you

    • The machine can take the first step. Starting a screen agent is easy.
    • Checking and deciding stay with you. Approving, watching, and declining stay your job.
    • This is what to practice. If the product safeguards drop, you hold the risk yourself.

    Not claimed The agent already works unsupervised in the products.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. Produkt-Autonomie ist durch Bestätigung und Watch-Mode by design nicht unbeaufsichtigt. Agents E2
    2. Mensch-Modell-Lücke auf Ausführungsevals: GAIA L3 Paper 0 %; WebArena Paper 14.41 % vs. 78 %; OSWorld Paper 12.24 % vs. 72 %. Agents E1
    Which rule (Operation)
    hitl-bleibt-hitl
    What follows (Consequence)
    Die gemessene Ausführungslücke und das Produktmodell bleiben Human-in-the-Loop, nicht unbeaufsichtigte Autonomie.
    How far it holds (Limit)
    „Autonom“ in Produkten ist durch Design nicht unbeaufsichtigt: confirmation, Watch mode, decline banking. Das Produktmodell ist Human-in-the-loop.
    When it tips (Tipping condition)
    In denselben geöffneten Produkt-Docs entfallen confirmation und Watch-Mode für denselben Computer-Use-Stack.
  3. F3 constructed (E0)

    Faster tasks are not the jobs picture

    Today

    Short writing tasks get measurably faster with AI. A human still checks. Firm use and the jobs picture are something else.

    What follows

    Tasks get faster. Jobs do not vanish for that reason.

    What this means for you

    • The machine can take the first step. Drafts and short texts arrive faster.
    • Checking and deciding stay with you. You check the text yourself.
    • This is what to practice. Even when more firms use AI, do not read task speed as the jobs picture.

    Not claimed Jobs are therefore permanently safe.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. Noy & Zhang Science 381: Writing-RCT n=453, Zeit −40 %, Qualität +18 %. Work and Education E1
    2. Census BTOS: Firmen-Adoption von KI 17–20 %. Work and Education E1
    Which rule (Operation)
    hitl-bleibt-hitl
    What follows (Consequence)
    Die gemessene Task-Beschleunigung bleibt eine HITL-Messung und trägt nicht die Beschäftigungsstatistik.
    How far it holds (Limit)
    Muster bis einschließlich BTOS 2026: HITL-Tasks beschleunigen in RCT/Feld; Beschäftigungsstatistik und Firmen-Adoption (17–20 %) bleiben hinter Frey 47 % und Eloundou 80 %.
    When it tips (Tipping condition)
    Census-BTOS-Firmenanteil an KI-Nutzung erreicht in derselben Berichtsreihe die Größenordnung 47–80 %.
  4. F4 constructed (E0)

    More computing draws more electricity

    Today

    Large data centres in the US no longer use a flat amount of electricity. Heavier tasks draw much more energy.

    What follows

    More computing draws more electricity. The flat phase is over.

    What this means for you

    • The machine can take the first step. Hitting send on a heavy request is cheap. The electricity is not.
    • Checking and deciding stay with you. Check plans against measured kilowatt-hours, not the story.
    • This is what to practice. Check the power link on the new measurement, not the old curve.

    Not claimed AI alone explains the whole rise versus cloud and crypto.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. US-DC-Strom: ~60 TWh 2014–16, 76 TWh 2018 (1.9 %), 176 TWh 2023 (4.4 %); GPU-accelerated Server sind der Bruch; E1 für 2014–2023-Best-Estimate. Energy and Chips E1
    2. Inferenzenergie hängt um Größenordnungen von der Aufgabenklasse ab (Luccioni et al.). Energy and Chips E1
    Which rule (Operation)
    compute-zieht-energie
    What follows (Consequence)
    Dieselbe Kopplung bleibt: mehr Compute und teurere Tasks ziehen gemessenen Datacenter-Strom, nicht nur Spec-TDP.
    How far it holds (Limit)
    LBNL: US-Ära der flachen Datacenter-kWh ist vorbei; Kausalanteil AI vs. Rest-Cloud vs. Crypto nicht fein aufgelöst. IEA ~415 TWh 2024 bleibt E2.
    When it tips (Tipping condition)
    LBNL-US-Datacenter-kWh kehren in einer neuen Best-Estimate-Reihe auf das flache Niveau der Ära vor dem GPU-Bruch zurück.
  5. F5 constructed (E0)

    The model filters. The lab checks.

    Today

    Models produce and filter guesses about proteins. Whether they hold is shown in the lab.

    What follows

    The model filters guesses. The lab remains the check.

    What this means for you

    • The machine can take the first step. Candidate lists arrive fast. They are not proof.
    • Checking and deciding stay with you. Binding and measuring stay lab work.
    • This is what to practice. Even if the lab gets more accurate, do not hand the check to the model score.

    Not claimed Everyone needs protein engineering as a career.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. RFdiffusion Binder-Nasslabor: 95 Designs × 5 Targets; overall experimental success rate 19 % (BLI); In-silico-Erfolg ersetzt Nasslabor nicht. Proteins E1
    2. Terwilliger: AF-Vorhersagen sind Hypothesen; Map–Model-Korrelation AF 0.56 vs. deposited 0.86. Proteins E1
    Which rule (Operation)
    hypothese-nicht-nasslabor
    What follows (Consequence)
    Struktur- und Designmodelle bleiben Hypothesenfilter; der Nasslabor-Schritt bleibt die gemessene Erfolgsinstanz.
    How far it holds (Limit)
    Vorhersage ist nicht Design. RFdiffusion Binder-Nasslabor overall experimental success rate 19 % (BLI); In-silico-Erfolg ersetzt Nasslabor nicht. AF-Vorhersagen sind Hypothesen.
    When it tips (Tipping condition)
    RFdiffusion overall experimental success rate (BLI, Paper-Protokoll) liegt bei ≥50 %.
  6. F6 constructed (E0)

    Lithium-ion runs at volume. Solid-state stays a pilot.

    Today

    At volume, electric cars run on lithium-ion cells. Solid-state appears in the reports as a pilot, not as a product in use.

    What follows

    Lithium-ion runs at volume. Solid-state remains a pilot.

    What this means for you

    • The machine can take the first step. Comparing cells by storage and price is ordinary series tech.
    • Checking and deciding stay with you. Do not count pilot lines as volume use.
    • This is what to practice. When solid-state shows up, tell a product measurement from a pilot claim.

    Not claimed Solid-state never arrives.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. IEA GEO 2026: EV-Batterie-Deployment 1.2 TWh in 2025; LFP über 55 % der EV-Batterien 2025. Batteries E2
    2. Frith/Lacey/Ulissi 2023: kommerzielle Zellen ≥270 Wh kg⁻¹ / ≥650 Wh L⁻¹; Zellpreis 101 $ kWh⁻¹ (2021); Zelle ≠ IEA-Pack. Batteries E2
    Which rule (Operation)
    mehr-desselben
    What follows (Consequence)
    Die eingesetzte Praxis bleibt Li-Ion auf TWh-Skala; Festkörper und Natrium bleiben außerhalb dieser Prämissen.
    How far it holds (Limit)
    Festkörper bleibt Herausforderung (Janek & Zeier 2023, E2); IEA: Vorteile in realen Anwendungen noch nicht gezeigt, all-SSB Prototyp-Stadium. Natrium bleibt Behauptung; 500 Wh/kg plus 1 000 Zyklen ist kein EV-Produkt.
    When it tips (Tipping condition)
    IEA-Berichte messen all-solid-state als eingesetztes EV-Produkt in Wh/kg-Pack-Skalierung neben Li-Ion-TWh, nicht nur als TRL/Pilot-Text.
  7. F8 constructed (E0)

    One hit is not enough

    Today

    One hit with a tool-using agent is not enough. Across many repeats, reliability breaks. The end state alone does not show that.

    What follows

    One hit is not enough. More tries show the same drop.

    What this means for you

    • The machine can take the first step. The answer appears quickly.
    • Checking and deciding stay with you. You still need it when the helper is gone.
    • This is what to practice. First alone, then with help, then alone again.

    Not claimed The public 2026 measurement is already open.

    Evidence and technical fields
    What already stands (Premises from the atlas)
    1. τ-bench function-calling pass^1: gpt-4o avg 48.2; pass^8 retail gpt-4o <25 %; Policy-Ablation Airline gpt-4o 33.2 → 10.8. Agents E1
    Which rule (Operation)
    mehr-desselben
    What follows (Consequence)
    Mehr Trials derselben Tool-Agent-Messung zeigen denselben Konsistenzbruch: pass^k liegt unter pass^1.
    How far it holds (Limit)
    τ-bench: User-Sim = LM; Reward = DB-Endzustand, nicht hinreichend für Policy-Treue. pass^k fällt steil; Live-Boards 2026 nicht geöffnet.
    When it tips (Tipping condition)
    τ-bench pass^8 liegt für dasselbe Modell und Domain nicht mehr unter pass^1 im Paper-Protokoll.

File as of 2026-09-14 · Source Bot B/C/D + Sprach-Bot Leserschicht 2026-09-14 · Review ungeprüft (Josef)