“COMPARISON OF DIAGNOSTIC AND TREATMENT RELIABILITY OF ARTIFICIAL INTELLIGENCE (AI) SOFTWARE WITH HUMAN EVALUATION USING PANORAMIC RADIOGRAPHS”

Authors

  • Chinta Brahmaja, MDS Author
  • Dr. Challagulla Anusha, MDS Author
  • A. Ratnaditya, MDS Author
  • Hari Pranathi Mallavarapu, MDS Author
  • Vyshnavi mididoddi, MDS Author
  • K. Aishwarya, MDS Author

DOI:

https://doi.org/10.4238/dpvmxq92

Keywords:

Artificial intelligence (AI); machine learning; digital dentistry, orthopantomography, pediatric radiograph, mixed dentition.

Abstract

Background/Aim: AI in pediatric dentistry is used for the purpose of making an accurate diagnosis and assisting clinicians, dentists, and pediatric dentists in clinical decision making, developing preventive strategies, and is also a time-saving procedure. The current state-of-the-art artificial intelligence-based natural language modelling has become so convincing that readers cannot tell the difference between human and machine-written texts. The present study was conducted to compare AI software with expert visual examination in diagnosing dental conditions in pediatric patients. Methodology: A total of 208 panoramic radiographs were collected from patients attending the Department of Pediatric and Preventive Dentistry at Meghna Institute of Dental Sciences. Each radiograph was evaluated independently by an expert team (Pediatric Dentist), who will provide detailed diagnostic reports. The same radiographs will also be analysed using an advanced AI software system. This generated a comprehensive report, including sensitivity and specificity metrics for comparison with expert evaluations. Results: The inter-examiner kappa values were consistently high across all diagnostic categories, ranging from 0.866 to 0.982, indicating excellent agreement among human examiners. AI also demonstrated good to excellent agreement, with kappa values ranging from 0.722 to 0.955. The highest agreement for AI was observed in the diagnosis of dental caries (κ = 0.955), closely approaching human agreement levels, while the lowest agreement was seen in periodontal pathology (κ = 0.722), reflecting moderate to good reliability. Conclusion: The present study compared the diagnostic performance of human examiners and an AI-based software system in detecting multiple dental pathologies on panoramic radiographs using expert consensus as a reference standard. Overall, human evaluators demonstrated consistently higher sensitivity across all assessed conditions, including dental caries, periodontal pathology, periapical lesions, root fragments or impacted tooth and prosthesis or missing teeth.

Downloads

Published

2026-09-23

Issue

Section

Articles