As a CISPA Helmholtz Center for Information Security spin-off, we are very happy and thankful about the long-lasting research collaboration. A special thanks to Prof. Jilles Vreeken who supports QuantPi since day 1. "Where Do Agents Differ? Interpretable Rule Discovery for Performance Differences Across Models and Data"
Aggregate benchmarks hide more than they reveal. We published a paper today at the ICML 2026 Workshop AI4GOOD (https://lnkd.in/ek7Mii98), together with Sascha Xu and Jilles Vreeken from CISPA Helmholtz Center for Information Security: "Where Do Agents Differ? Interpretable Rule Discovery for Performance Differences Across Models and Data" The core finding: when you compare two agentic backends on the same benchmark, the aggregate accuracy gap tells you almost nothing about where the difference actually lives. On SWE-bench, we compared five OpenHands backends. Claude Opus and GPT-5 are separated by about 6 points overall. Unremarkable. But within long-trajectory tasks - the kind where an agent has to sustain coherent reasoning across many steps - Opus resolves 90% and GPT-5 resolves 56%. A 34-point gap that the aggregate number buries completely. And the reverse exists too: a narrow task regime where GPT-5 reaches 89% and Opus drops to 36%. This is not a corner case. Across all ten backend pairs we tested, seven contained at least one regime where the weaker backend outperformed the stronger one. What we see with enterprise customers mirrors this exactly. Model selection decisions are being made on aggregate scores over benchmark distributions that don’t reflect the actual task distribution in production. The result is systems that perform well on paper and fail on the specific slice of work that matters. The method we apply - differential subgroup discovery - finds these structured differences automatically and describes them in interpretable rules. Not “this model is better.” But “this model is better for tasks with these specific characteristics. That distinction is what makes evaluation actually useful. Last but not least, HUGE thanks to Sascha Xu and Jilles Vreeken for their impressive work on this.