FinResearchBench II introduces a scalable pipeline for generating evaluation rubrics for financial deep research reports without human experts in the final loop. Built from 104 real-world queries, the benchmark synthesizes 14,450 candidate rubrics and retains 2,600 consensus-derived gold rubrics through consistency and distinguishability filters. LLM-based evaluation achieved 98.67% agreement with human experts on unanimous items, enabling differentiated rankings across 10 deep research systems with pass rates from 58.58% to 22.23%.
No score is assigned. Sources and their independence are shown in the citation chain below.