OE AI tool outperforms general-purpose models in clinical accuracy and utility

For the past two weeks, our independent team of statisticians, AI evaluation experts, clinical AI researchers, and clinicians was given a unique opportunity to test one question: “How well do different AI tools answer user questions on the OpenEvidence (OE) platform?” We were given access to 620 benchmark questions sampled from real queries submitted to OE, and 149 practicing physicians across 36 states made blinded head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT5.5) and OE’s specialized clinical AI tool. Graders were matched to questions that matched their specialty to maximize evaluation accuracy. Today we are sharing the results on arXiv and the benchmark dataset on HuggingFace. Following our prespecified statistical analysis, OE scored highest on all five dimensions of accuracy, clinical utility, source quality, verifiability, and completeness. On the primary endpoint of win differences (a model’s win rate minus its loss rate in head-to-head comparisons), OE led by 25–39 percentage points all these dimensions (p < 0.001). Results were consistent across our *many* sensitivity analyses. So, how do these results square with recent work reporting general-purpose LLMs outperform OE? We think both findings can be true: when a specialized tool is evaluated in a setting outside of its intended workflow, it may underperform. But the more relevant question for AI application developers is whether their tool is delivering value in its intended use case for its target users. On that measure, our findings are positive. At least for now, the hard work of engineering and customizing an AI system can still pay off, delivering meaningful performance gains to its target users. Please see the paper for important details and nuances of the work. Our hope is that this starts a healthy discussion on how to best evaluate AI systems, because no evaluation is perfect and every evaluation involves tradeoffs. 📎 : https://lnkd.in/eJj9sVpZ 🤗 : https://lnkd.in/ewBxmbnT 💻: https://lnkd.in/e_Mv4F_x The incredible team behind this work: Vishal Patel, Patrick Heagerty, Yifan Mai, Venkatesh S., Patrick Vossler, Jialin Ouyang, Bapu Jena

  • chart

Thank you, very helpful write up for people working on LLMs or LLM agents in the clinical decision support space. This confirms what we are seeing from from users of specialized AI applications. So, really good to see evidence accumulating that high effort harness development and prompt engineering can pay off for specialized used cases, as you say for now.

And context is so important!! Love this work Jean!

See more comments

To view or add a comment, sign in

Explore content categories