AMBOSS Newsroom
Product Update

Independent benchmark: AMBOSS ranks #1 again, outperforming OpenEvidence, Doximity, and Glass Health

Published on
July 28, 2026
AMBOSS Newsroom
Product Update

Independent benchmark: AMBOSS ranks #1 again, outperforming OpenEvidence, Doximity, and Glass Health

Published on
July 28, 2026
Contributors
AMBOSS empowers students, medical professionals, and educators with trusted knowledge and smart solutions.
Read about our privacy policy.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

A fully revised version of the independent Stanford–Harvard benchmark NOHARM (Numerous Options Harm Assessment for Risk in Medicine) was released in July this year. The study evaluated clinical AI using real consultation cases with more than 4,000 expert-reviewed clinical decisions. AMBOSS AI LiSA remains the most accurate and complete clinical AI in the benchmark, surpassing every other clinical AI tested, including OpenEvidence, Doximity Ask, and Glass Health.

The study also found that clinicians who use high-quality AI support effectively outperform clinicians who don’t use AI. The implication is clear: Patient safety depends not only on which clinical AI tool is used, but also on whether one is used at all.

Measuring what matters most in clinical AI

NOHARM is a clinical-safety benchmark based on real physician-to-specialist eConsults from Stanford Health Care. Each case reflects the uncertainty, missing context, and ambiguity of real clinical practice.

The scale of data used is substantial:

  • 100 real eConsult cases with 1,000 modified versions, designed to test the consistency of a model when a case is presented differently
  • 4,249 management options rated by 29 board-certified physicians, producing 12,747 expert annotations of appropriateness and harm severity

The study was conducted by ARISE (AI Research and Science Evaluation), a Stanford–Harvard network established in 2024 that brings together clinicians and data scientists to independently evaluate AI in health care. 

Every AI response was scored against expert rubrics for two kinds of error:

  • Errors of commission: recommending something inappropriate
  • Errors of omission: leaving out something important

Both were weighted by how much harm the error could cause (mild, moderate, or severe) and combined into a single severity-weighted F1 score. The higher the score, the more accurate and complete the AI recommendations.

Clinicians and AI complement each other when AI recommendations are heeded

Purpose-built clinical AI outperformed general-purpose models, with AMBOSS AI leading the field. On the benchmark's primary measure, the severity-weighted F1 score, which combines correctness and completeness, AMBOSS AI LiSA scored highest at 86.15%. 

AMBOSS AI also had the lowest severe-error rate of any system tested, with potential for serious harm in just 2.9% of cases. For perspective, the benchmark also included a "Do Nothing" baseline that carried severe-harm potential in 37% of cases: A reminder that inaction is rarely the safest option, and that the honest question about AI safety is always "compared with what?"


The randomized study of 101 attending physicians showed AI assistance improved their management plans. But it also revealed a blind spot: Clinicians frequently ignored valuable AI recommendations. Had physicians used AI properly and applied appropriate recommendations, they would have improved their performance.

What this means for clinicians, educators, and health systems

NOHARM points to a clear path forward and a shared responsibility.

Choose the safest tool available and demand independent evidence. Not all AI is built for clinical care; any compromise on safety or completeness puts patients at risk. Clinicians and institutions should demand independent, rigorous, real-world safety benchmarks before adopting any clinical AI tool.

Teach and practice responsible use of AI. AI is only useful when clinicians know how to use it appropriately. As it becomes part of everyday practice, teaching students and residents to use it responsibly is essential to guard against overreliance, deskilling, and failing to develop core clinical skills. 

At AMBOSS, these are the standards we uphold. Ranking first in an independent safety benchmark reflects the work of hundreds of experienced physicians and engineers and the millennia-old commitment: first, do no harm.

Review the findings: NOHARM preprint on arXiv 

Explore the leaderboard: ARISE interactive results