Despite passing medical licensing exams, large language models are not yet safe for autonomous clinical decision-making because they optimize for the most probable text rather than the critical task of identifying rare, high-consequence diagnoses, according to a new perspective published on arXiv. The core problem lies in information gathering under uncertainty: LLMs fail to exhibit essential triage behaviors like broadening differential diagnoses, seeking missing red flags, and properly escalating concern when dangerous diagnoses cannot be excluded.
Why it matters: As healthcare organizations increasingly deploy LLMs for patient triage and symptom assessment, this research highlights critical safety gaps that could lead to missed diagnoses and patient harm if autonomous systems lack proper clinical oversight.