3 min. read

Is AI Actually “Reasoning” in Medicine, or Just Repeating What It Memorized?

  • Healthcare AI
  • Thought Leadership

One of the biggest questions hanging over AI in healthcare right now isn’t whether large language models (LLMs) like ChatGPT can pass a medical exam. They can. It’s whether they’re actually reasoning through a case, or just recalling patterns from the mountains of medical literature they were trained on. That distinction matters enormously: a model that’s memorized textbook cases might fall apart the moment it sees something genuinely new.

Dr. Sheppert and colleagues designed a clever way to test this directly, in a new study published in the Journal of the American Medical Informatics Association (JAMIA).

The Setup

The team pulled 2,000 real clinical case reports from PubMed Central and split them into two groups:

  • 1,000 “contaminated” cases from 2021–2022, published early enough that they were almost certainly part of the data these AI models were trained on.
  • 1,000 “clean” cases from 2025, published after the models’ training data was locked in, meaning the models could never have seen them before.

 

If an LLM does better on the older cases, that’s a red flag: it suggests the model is leaning on memorized answers rather than genuine clinical reasoning. If performance holds steady across both groups, that’s evidence the model is actually working through the case logic in real time.

What They Found

The results were strikingly consistent. Diagnostic accuracy was nearly identical between the two groups (about 67% either way), and detailed text-similarity analysis showed no meaningful sign that the models were echoing memorized language from older cases. In other words, the models performed just as well on cases they couldn’t have memorized as on ones they theoretically could have.

Why This Matters

For clinicians, health systems, and patients alike, this is a genuinely reassuring data point in an area full of hype and skepticism in both directions. It doesn’t mean AI diagnostic tools are ready to replace clinical judgment: accuracy in the high-60% range still leaves plenty of room for error, and this study focused on written case reports rather than the messiness of real-time patient care. But it does push back on a common and reasonable worry: that these tools look smart mainly because they’ve seen the answer key before.

As AI tools continue to make their way into clinical workflows, careful, methodical studies like this one are exactly what’s needed to separate genuine capability from statistical sleight of hand.

Read the full study: Sheppert AP, Adams B, Sheppert AD, Riley S. “Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine.” Journal of the American Medical Informatics Association, 2026. https://doi.org/10.1093/jamia/ocag069

This post is intended for general educational purposes and does not constitute medical advice.