
- Connected Care
- Healthcare AI
- Physician Burnout
- Clinical Workflow Automation
3 min. read

One of the biggest questions hanging over AI in healthcare right now isn’t whether large language models (LLMs) like ChatGPT can pass a medical exam. They can. It’s whether they’re actually reasoning through a case, or just recalling patterns from the mountains of medical literature they were trained on. That distinction matters enormously: a model that’s memorized textbook cases might fall apart the moment it sees something genuinely new.
Dr. Sheppert and colleagues designed a clever way to test this directly, in a new study published in the Journal of the American Medical Informatics Association (JAMIA).
The team pulled 2,000 real clinical case reports from PubMed Central and split them into two groups:
If an LLM does better on the older cases, that’s a red flag: it suggests the model is leaning on memorized answers rather than genuine clinical reasoning. If performance holds steady across both groups, that’s evidence the model is actually working through the case logic in real time.
The results were strikingly consistent. Diagnostic accuracy was nearly identical between the two groups (about 67% either way), and detailed text-similarity analysis showed no meaningful sign that the models were echoing memorized language from older cases. In other words, the models performed just as well on cases they couldn’t have memorized as on ones they theoretically could have.
For clinicians, health systems, and patients alike, this is a genuinely reassuring data point in an area full of hype and skepticism in both directions. It doesn’t mean AI diagnostic tools are ready to replace clinical judgment: accuracy in the high-60% range still leaves plenty of room for error, and this study focused on written case reports rather than the messiness of real-time patient care. But it does push back on a common and reasonable worry: that these tools look smart mainly because they’ve seen the answer key before.
As AI tools continue to make their way into clinical workflows, careful, methodical studies like this one are exactly what’s needed to separate genuine capability from statistical sleight of hand.
Read the full study: Sheppert AP, Adams B, Sheppert AD, Riley S. “Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine.” Journal of the American Medical Informatics Association, 2026. https://doi.org/10.1093/jamia/ocag069
This post is intended for general educational purposes and does not constitute medical advice.