The work setup makes the difference
Today
When you measure coding agents, you often measure the setup around the model, not the model alone. Many runs fail mid-way on code changes.
What follows
The result comes from the setup, not from the bare model.
What this means for you
- The machine can take the first step. Starting small code changes is easier with the helper setup.
- Checking and deciding stay with you. You check failed mid-steps, not only the final result.
- This is what to practice. Set the result with the setup beside the result without it.
Not claimed The setup makes code changes reliable.
Evidence and technical fields
- What already stands (Premises from the atlas)
- Which rule (Operation)
- scaffold-nicht-modell
- What follows (Consequence)
- Dieselbe Ausführungsmessung bleibt an ACI/Spec-Text und Scaffold gebunden, nicht an das nackte Modell.
- How far it holds (Limit)
- SWE-agent ACI vs. Shell-only Lite 18.00 % vs. 11.00 %; Scaffold-Inflation: Spec-Text und ACI bewegen Scores stärker als das Modell allein. Autonomie-Scores sind Scaffold-Scores.
- When it tips (Tipping condition)
- Auf demselben SWE-agent-Lite- oder SWE-Pro-Ablation-Protokoll liegt der Score ohne ACI/Augmentation nicht mehr unter dem Scaffold-Lauf.