IBM Research and Hugging Face introduced VAKRA, a tool-grounded, executable benchmark evaluating how AI agents reason and act in enterprise-like environments. VAKRA measures compositional reasoning across APIs and documents, featuring over 8,000 locally hosted APIs spanning 62 domains with tasks requiring 3-7 step reasoning chains. The benchmark includes four capability categories, including API chaining and tool selection. Developers Ankita Naik, Danish, Ben, Anupama Murthi, Praveen Venkateswaran, Siyu, and Ayhan Sebin report that models perform poorly on VAKRA.
No score is assigned. Sources and their independence are shown in the citation chain below.