Please login or create an account. If you do not have access to this content, you will be shown a 30 second preview and licensing options.

  • Presentation

Limitations of AI in Medicine: Reasoning, Hallucinations, Patient Safety, and Real-World Evidence

Description

The speaker, a dermatologist and professor, argues that while AI has impressive medical capabilities, it still has major limitations that clinicians must understand. He highlights four main problems: AI’s illusion of reasoning, hallucinations, risks when patients use AI on their own, and the lack of strong real-world evidence. Studies show that large language models often rely on pattern matching rather than true clinical reasoning, perform poorly on tasks that require updating judgments with new information, and rarely admit when they cannot answer a question, instead confidently fabricating responses. In medicine, hallucinations can lead to false radiology findings, excessive drug interaction alerts, and treatment recommendations that conflict with guidelines. Retrieval-augmented systems help but still frequently produce unsupported claims. He also warns that patients can misinterpret or misuse AI advice, with studies showing non-experts struggle to judge quality, may trust low-accuracy AI as much as physicians, and can receive opposite triage advice based on subtle prompt differences. Finally, he notes that most evaluations focus on medical knowledge rather than practical workflow tasks, and while AI scribes may improve satisfaction, productivity gains remain modest. The overall message is that AI is promising but not yet reliable enough to replace critical clinical judgment, and clinicians must stay involved to ensure safety and usefulness.

View more

Conclusions

  • Current medical LLMs often appear to reason but are largely pattern-matching and can lose accuracy when familiar answer structures are removed.
  • These models still struggle to update their clinical judgments appropriately when presented with new information, often overreacting and failing to match expert reasoning.
  • LLMs show weak metacognition: they rarely recognize when a question is impossible to answer and tend to respond with unwarranted confidence.
  • Hallucinations remain a major problem, with models generating plausible-sounding but false medical statements, recommendations, and even fabricated citations.
  • Retrieval-augmented generation improves grounding but does not eliminate unsupported claims, so citations still require careful verification.
  • Because medical knowledge changes quickly and requires precision and context, healthcare is especially vulnerable to model obsolescence and cascading errors.
  • Patients are at particular risk because they may trust low-quality AI advice and often cannot extract correct answers from chatbot interactions.
  • The same model can give opposite advice to similar patients depending on wording, showing that prompt phrasing can strongly and unpredictably change outputs.
  • Consumer health chatbots still perform poorly at triage, especially at the clinical extremes where overtriage and undertriage can have serious consequences.
  • Most current LLM evaluations focus on exams and knowledge questions rather than real clinical workflows or patient data, limiting what the evidence can support.
  • In EHR settings, LLMs are more reliable at retrieving information than taking actions such as ordering, documenting, or referring.
  • AI scribes can improve clinician satisfaction and save some time, but the productivity gains so far are modest rather than transformative.
  • Overall, AI is promising in medicine but is not yet reliable enough to be used uncritically, so clinicians need to remain engaged in how these tools are developed and deployed.
  • Bedi S et al. Fidelity of Medical Reasoning in Large Language Models. JAMA Netw Open. 2025 Aug 1;8(8):e2526021.#10.1001/jamanetworkopen.2025.26021
  • McCoy LG et al. Assessment of large language models in clinical reasoning: a novel benchmarking study. NEJM AI. 2025;2(10).#10.1056/aidbp2500120
  • Griot M et al. Large Language Models lack essential metacognition for reliable medical reasoning. Nat Commun. 2025 Jan 14;16(1):642.#10.1038/s41467-024-55628-6
  • Kim, Y., et al., "Medical hallucinations in foundation models and their impact on healthcare," arXiv preprint arXiv:2503.05777 (2025).
  • Wu, K et al. An automated framework for assessing how well LLMs cite relevant medical references. Nat Commun 16, 3615 (2025).#10.1038/s41467-025-58551-6
  • Shekar S et al. People overtrust AI-generated medical advice despite low accuracy. NEJM AI. 2025;2(6).#10.1056/aioa2300015
  • Bean AM et al. Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nat Med. 2026 Feb;32(2):609-615.#10.1038/s41591-025-04074-y
  • Ramaswamy et al., Nature Medicine (2026).
  • Bedi S et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025 Jan 28;333(4):319-328.#10.1001/jama.2024.21700
  • Jiang Y et al. MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents. NEJM AI (2025).#10.1056/aidbp2500144