Claude Opus 5 achieved the highest pass rate, 12.3%, on ATLAS-Finance, a new benchmark of 100 expert-level tasks set in 13 realistic financial firm environments. Claude Fable 5.1 and GPT-6 Astra scored 12.0% and 11.3%, while the other eight models tested fell below 10%. Common failures included applying wrong financial logic, omitting required scope, and failing to propagate correctly calculated values downstream. Tasks take human experts 15-30 hours and are graded against expert-authored rubrics.
No score is assigned. Sources and their independence are shown in the citation chain below.