What this paper finds
- A published six-agent reservoir history-matching system improved error by 95% on a toy benchmark, 69% on a moderate one, and 13% on a real field.
- Across seventeen frontier models in four agent configurations, the best reached 59.5% on paired tasks requiring both acting and abstaining correctly — and abstention was found largely independent of task-solving capability.
- Agents were observed executing irreversible actions before recognising they should have stopped.