Beyond the Hype: Large Language Models, Clinical Reasoning, and the Future of Emergency Triage
Journal of Practical Emergency Medicine,
Vol. 13 No. 1 (2026),
5 June 2026,
Page e1
https://doi.org/10.22037/jpem.v13i1.48754
Large language models (LLMs) are moving into clinical medicine faster than the evidence needed to define their safest role. Emergency medicine makes that gap particularly difficult to ignore. In the emergency department (ED), a convincing answer is not necessarily a safe answer, and a seemingly small error in prioritization may delay care for a patient with a time-sensitive condition. The relevant question, therefore, is no longer whether an LLM can produce a plausible diagnosis or assign an acuity category. It is whether the use of such a model can improve emergency care without introducing an unacceptable risk of under-triage, misplaced confidence, or loss of clinically important context.
Triage is an especially demanding test case. Decisions are made early, often with incomplete information, before the diagnostic picture has matured. A patient may be unable to provide a coherent history; laboratory and imaging results may not yet exist; and clinical deterioration can occur after the initial assessment. Moreover, triage errors are asymmetric. Over-triage consumes resources, but under-triage may postpone recognition of stroke, sepsis, acute coronary syndrome, internal bleeding, or another immediately threatening condition. Any technology positioned at this point in the care pathway should therefore be judged not only by overall accuracy but also by the clinical consequences of its errors.
Early evidence has nevertheless been impressive enough to justify serious attention. Williams et al. evaluated an LLM using data from 251,401 adult ED visits. In a balanced sample of 10,000 pairs of patients with different Emergency Severity Index (ESI) levels, the model correctly identified the patient with higher acuity in 8,940 pairs, corresponding to an accuracy of 89%. In a 500-pair subset reviewed by physicians, LLM accuracy was 88% compared with 86% for physicians (1).
These results are encouraging, but the task should not be mistaken for prospective triage. The model was asked to compare two written clinical histories and determine which represented the higher-acuity patient. It was not independently assessing patients arriving at a triage desk, nor was it operating prospectively within an ED workflow. The study therefore demonstrated that an LLM could extract useful acuity information from clinical text; it did not establish that autonomous LLM-based triage is safe.
Other evidence makes that distinction even more important. Zaboli et al. retrospectively compared ChatGPT-4.0 with human triage decisions in 2,658 ED patients. Agreement between human and AI triage was low, with a Cohen κ of 0.125 (95% CI, 0.100–0.134). Human triage also performed better in predicting clinically important outcomes. For 30-day mortality, the area under the receiver operating characteristic curve was 0.88 for human triage and 0.70 for AI triage (P<0.001). For the need for life-saving interventions, the corresponding values were 0.98 and 0.87 (P=0.014). The authors specifically identified lower sensitivity in high-risk patients and consequent under-triage as a concern (2).
Yet another prospective study produced a different picture. Arslan et al. compared ChatGPT Plus, Copilot Pro, and triage nurses against ESI assessments assigned by an emergency physician. Overall accuracy was 66.5% for ChatGPT, 61.8% for Copilot, and 65.2% for nurses. Notably, identification of high-acuity patients was more frequent with ChatGPT and Copilot than with nurses in that study 87.8% and 85.7% versus 32.7%, respectively (3).
The apparent contradiction among these studies is itself informative. LLM performance cannot be treated as a fixed property of “AI.” Results depend on the model and version tested, the information supplied to it, the prompt, the reference standard, the triage system, the patient population, and the exact task being evaluated. A model that performs well when ranking two written cases may not perform equally well when assigning an absolute triage category. Likewise, good average accuracy can conceal clinically unacceptable errors in a small but high-risk subgroup.
How information is supplied to the model may be as important as the model itself. Yazaki et al. examined a retrieval-augmented generation (RAG) approach using 100 simulated triage scenarios derived from modified cases from the Japanese National Examination for Emergency Medical Technicians. GPT-3.5 combined with RAG achieved 70% correct triage, compared with 35% and 38% for two emergency medical technicians and 50% and 47% for two emergency physicians. Under-triage occurred in 8% of cases with GPT-3.5 plus RAG, compared with 33% for GPT-3.5 without RAG and 39% for GPT-4 without RAG (4). The findings suggest that connecting an LLM to task-specific knowledge may improve performance. At the same time, the study used simulated scenarios rather than live emergency encounters, which limits how far these results can be extended to routine practice.
Gaber et al. similarly evaluated several LLM configurations and a RAG-assisted workflow using 2,000 cases derived from the MIMIC-IV ED database. Their framework addressed triage, specialty referral, and diagnosis rather than treating the model as a stand-alone question-answering system (5). This is an important shift in emphasis. The relevant unit of evaluation may ultimately be the LLM-enabled clinical workflow, not the language model in isolation.
There are also reasons to remain cautious about extrapolating high performance on constrained tasks to autonomous clinical decision-making. Hager et al. evaluated selected LLMs using a curated dataset of 2,400 real patient cases involving four common abdominal pathologies. The tested models performed worse than physicians across the evaluated pathologies and showed limitations in following diagnostic and treatment guidelines, interpreting laboratory results, and responding consistently to changes in instructions (6). These findings should not be generalized automatically to every newer model—the technology evolves rapidly—but they demonstrate why conventional medical benchmarks are an incomplete proxy for bedside reliability.
The same caution applies to the concept of the AI “co-pilot.” Human supervision alone does not guarantee that an LLM will improve decisions. In a randomized clinical trial involving 50 physicians from internal medicine, family medicine, and emergency medicine, Goh et al. compared physicians using GPT-4 with physicians using conventional resources. Median diagnostic reasoning scores were 76% and 74%, respectively, an adjusted difference of 2 percentage points (95% CI, −4 to 8; P=0.60). Interestingly, the LLM operating alone scored 16 percentage points higher than the conventional-resources physician group (95% CI, 2–30; P=0.03) (7). Only five participants were emergency physicians, so the results are not an ED-specific effectiveness trial. Nevertheless, they expose an important problem: access to a capable model does not automatically translate into better human performance. Interface design, clinician training, appropriate reliance, and the way disagreement between clinician and model is handled may determine whether assistance is useful.
For this reason, the first successful applications of LLMs in emergency care may emerge from narrower tasks in which benefit can be measured without transferring final clinical authority to the model. Song et al. provide an instructive example. In a study involving six emergency physicians and 50 representative ED cases, LLM-assisted discharge documentation reduced median writing time from 69.5 seconds (95% CI, 65.5–78.0) to 32.0 seconds (95% CI, 29.5–36.0; P<0.001). LLM-assisted notes also received higher ratings than manually written notes for completeness, correctness, conciseness, and clinical utility (all P<0.001) (8). The study was performed in a single tertiary center and involved a limited validation set, but it illustrates a practical principle: the safest path to clinical adoption may begin with bounded tasks that reduce workload while leaving consequential decisions under direct clinician control.
Emergency triage requires an equally pragmatic standard. Before an LLM is used as more than supervised decision support, prospective multicenter studies should evaluate outcomes that matter to patients rather than relying predominantly on agreement with historical labels. Under-triage deserves particular attention, as do delays to time-sensitive treatment, unexpected intensive care unit admission, early ED return, and adverse clinical outcomes. Performance should also be examined across patient groups and clinical presentations rather than reported only as a single aggregate accuracy measure.
Model identity and version should be treated as part of the intervention. An evaluation performed with one version cannot be assumed to remain valid after a major model update. Similarly, changes in prompts, retrieval sources, or local workflow may alter performance. Clinical deployment will therefore require version control, post-implementation surveillance, and an auditable record of when AI-generated recommendations are accepted or overridden.
Accountability must be designed before deployment rather than after an adverse event. If an LLM assigns a lower acuity level than the triage clinician, should its recommendation be ignored, reviewed, or trigger a mandatory second assessment? If the model recommends escalation and the clinician disagrees, what should be documented? These are not peripheral legal questions. They determine how the technology will influence actual behavior in the ED.
The current evidence does not support either extreme position. It would be premature to dismiss LLMs as unreliable text generators when they have demonstrated substantial ability to interpret clinical narratives and, in selected tasks, match or exceed human performance. It would be equally premature to treat those findings as evidence that autonomous AI triage is ready for routine deployment. What has been established is capability under specific experimental conditions. What remains to be established is dependable clinical benefit.
For emergency medicine, that distinction matters. The front door of the hospital is an unforgiving place to discover that benchmark performance does not translate into patient safety. For now, the most defensible role for LLMs in emergency triage is not that of an autonomous gatekeeper, but of a supervised clinical co-pilot whose recommendations can be questioned, overridden, and continuously evaluated. The threshold for moving beyond that role should be set by prospective evidence of safer or better care not by the fluency of the model's answers.