3 min. read

When AI Jumps to Conclusions Faster Than Doctors Do

  • AI Healthcare
  • Clinical Intelligence Platform
  • Healthcare AI

Anchoring bias is one of the oldest and best-documented traps in medical decision-making: a patient mentions one detail (say, a recent trip, or a family member’s illness) and that detail can quietly hijack a clinician’s differential diagnosis, even when it isn’t actually the most likely explanation. Physicians are trained to recognize and guard against this. The open question is whether AI models, increasingly used as diagnostic aids, are just as susceptible, or worse still.

A study co-authored by Dr. Alexander Sheppert, DO, PhD, MBA published in the International Journal of Medical Informatics, put that question to a direct, head-to-head test.

The Setup

The researchers built nine pairs of internal medicine outpatient vignettes. Each pair had two versions:

  • An anchored version, where the patient casually mentions something suggestive of a plausible-but-unlikely diagnosis.
  • A matched control version, identical in every way except that suggestive detail was removed.

Twenty internal medicine residents and five attending physicians each ranked their top five differential diagnoses for these vignettes (225 responses in total). Then eight different large language models were run through the exact same forced-choice exercise, so the comparison was truly apples-to-apples.

What They Found

The AI models anchored on the misleading detail far more readily than the human physicians did. Statistically, the models were about four times more likely than resident physicians, and about four times more likely than attending physicians, to rank the suggested (but less likely) diagnosis at the top of their list, once the anchoring detail was present.

Put simply: in this controlled setting, physicians resisted the bait better than the AI did.

Why This Matters

This finding lands as an important counterweight to the more optimistic narrative around AI in diagnosis. It’s not that these models are unsophisticated. They can hold their own on many complex reasoning tasks. But bias isn’t the same thing as accuracy, and this study suggests LLMs may be particularly vulnerable to a specific, well-known human cognitive trap, in some cases more so than the physicians they’re meant to support.

The authors are careful to note that these were structured, forced-choice vignettes, not the messier, open-ended reality of an actual patient encounter, so they call for further real-world testing before drawing firm conclusions about clinical deployment. But the takeaway for now is a useful one: as AI tools get folded into clinical workflows, they need to be evaluated not just on raw accuracy, but on how they get to their answers and where they might quietly go wrong.

Read the full study: “Large language models exhibit greater diagnostic anchoring than physicians in a forced-choice vignette study.” International Journal of Medical Informatics, Vol 219, 2026. https://doi.org/10.1016/j.ijmedinf.2026.106550

This post is intended for general educational purposes and does not constitute medical advice.