ARTIFICIAL INTELLIGENCE VS. HUMAN TRIAGE: A SYSTEMATIC REVIEW OF DIAGNOSTIC ACCURACY AND CLINICAL OUTCOMES IN THE EMERGENCY DEPARTMENT: A SYSTEMATIC REVIEW

Authors

  • Hassan Abdullah Alasmari Author
  • Kholoud Khalid Almowald Author
  • Ibrahim Salem Alfaifi Author
  • Wejdan Hussain Alqatifi Author
  • Hashim Haider Alawam Author
  • Lena Mohammed Almutairi Author
  • Mohammed Fahad Salem Bajaber Author
  • Sari Salem Alqahtani Author
  • Abdulrahman Ali Almurshed Author
  • Sara Jamil Nasser Author

DOI:

https://doi.org/10.4238/ha9w4048

Keywords:

Artificial Intelligence, Human Triage, Diagnostic Accuracy, Clinical Outcomes, Emergency Department, Systematic Review.

Abstract

Objective: To systematically evaluate diagnostic accuracy and clinical outcomes of AI-assisted versus human triage in the emergency department.

Methods: PRISMA 2020-compliant systematic review was applied. Six databases (PubMed/MEDLINE, Embase, Scopus, CINAHL, Cochrane CENTRAL, IEEE Xplore) were searched through March 31, 2025. Eligible studies directly compared AI triage systems (machine learning [ML], deep learning [DL], natural language processing [NLP], or large language models [LLMs]) against validated human triage instruments (ESI, MTS, KTAS, or equivalent). Primary outcomes were diagnostic accuracy metrics (AUROC, sensitivity, specificity); secondary outcomes included mortality, hospital admission, ICU admission, and ED length of stay. Quality was assessed using QUADAS-2 and ROBIS.

Results: Eleven studies (>12.6 million ED encounters; 2018–2025) were included. ML and DL models consistently outperformed human triage instruments (AUROC: 0.85–0.94 vs. 0.69–0.80), with simultaneous reductions in undertriage and overtriage. Prospective implementation studies demonstrated improved patient flow and reduced triage inequity across racial and ethnic groups. LLMs underperformed trained triage professionals in all comparisons (κ = 0.54–0.67) and exhibited systematic overtriage.

Conclusion: ML and DL triage systems demonstrate clinically meaningful accuracy advantages over human triage. LLMs are not ready for standalone deployment. Evidence best supports an augmentative model in which AI enriches rather than replaces human triage judgment.

Downloads

Published

2026-06-25

Issue

Section

Articles