A ReAct agent using GPT-4.1 achieved 77.4% average success on AppWorld across five runs but completed all five runs for only 53.0% of tasks, a 24.4-point consistency gap. The authors introduce the Consistency Analyzer, which resamples recorded trajectories to identify flip-prone decisions, and consistency guidelines that halve the gap to 12.0 points without reducing average accuracy. They argue standard Mean@k benchmarks hide this unreliability, unlike Pass^k, which requires success on every run.
No score is assigned. Sources and their independence are shown in the citation chain below.