What this paper finds
- On an engineering benchmark spanning nine domains, the best of twenty-seven models reached 65.4% final-answer accuracy but only 42.7% reasoning-trace fidelity — and 79.5% of frontier-model errors were arithmetic rather than conceptual.
- A published six-agent system improved reservoir history matching by 95% on a toy case, 69% on a moderate one, and 13% on a real field. Performance collapses precisely as problems become real.
- In a 10,000-trial peer-reviewed study, equipping a model with task-specific deterministic tools moved accuracy from 11% to 84% and from 36% to 95% — outperforming both retrieval augmentation and a general code interpreter.