ARTIFICIAL INTELLIGENCE VS. HUMAN TRIAGE: A SYSTEMATIC REVIEW OF DIAGNOSTIC ACCURACY AND CLINICAL OUTCOMES IN THE EMERGENCY DEPARTMENT: A SYSTEMATIC REVIEW
DOI:
https://doi.org/10.4238/ha9w4048Keywords:
Artificial Intelligence, Human Triage, Diagnostic Accuracy, Clinical Outcomes, Emergency Department, Systematic Review.Abstract
Objective: To systematically evaluate diagnostic accuracy and clinical outcomes of AI-assisted versus human triage in the emergency department.
Methods: PRISMA 2020-compliant systematic review was applied. Six databases (PubMed/MEDLINE, Embase, Scopus, CINAHL, Cochrane CENTRAL, IEEE Xplore) were searched through March 31, 2025. Eligible studies directly compared AI triage systems (machine learning [ML], deep learning [DL], natural language processing [NLP], or large language models [LLMs]) against validated human triage instruments (ESI, MTS, KTAS, or equivalent). Primary outcomes were diagnostic accuracy metrics (AUROC, sensitivity, specificity); secondary outcomes included mortality, hospital admission, ICU admission, and ED length of stay. Quality was assessed using QUADAS-2 and ROBIS.
Results: Eleven studies (>12.6 million ED encounters; 2018–2025) were included. ML and DL models consistently outperformed human triage instruments (AUROC: 0.85–0.94 vs. 0.69–0.80), with simultaneous reductions in undertriage and overtriage. Prospective implementation studies demonstrated improved patient flow and reduced triage inequity across racial and ethnic groups. LLMs underperformed trained triage professionals in all comparisons (κ = 0.54–0.67) and exhibited systematic overtriage.
Conclusion: ML and DL triage systems demonstrate clinically meaningful accuracy advantages over human triage. LLMs are not ready for standalone deployment. Evidence best supports an augmentative model in which AI enriches rather than replaces human triage judgment.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

