Article Data

  • Views 172
  • Dowloads 117

Original Research

Open Access

Can AI see what clinicians see? A comparative study of GPT-5 and Gemini 2.5 in radiographic evaluation of regenerative endodontic treatments in immature permanent teeth

  • Enes Mustafa Aşar1,*,
  • Murat Selim Botsali1

1Department of Pediatric Dentistry, Selcuk University, 42130 Konya, Turkey

DOI: 10.22514/jocpd.2026.115 Vol.50,Issue 5,September 2026 pp.53-70

Submitted: 11 March 2026 Accepted: 14 April 2026

Published: 03 September 2026

*Corresponding Author(s): Enes Mustafa Aşar E-mail: enesmustafa.asar@selcuk.edu.tr

Abstract

Background: This study evaluates the diagnostic performance and agreement of GPT-5 and Gemini 2.5 in interpreting regenerative endodontic treatments (RET) on periapical radiographs compared with an expert reference standard. Methods: This retrospective diagnostic accuracy study included 51 paired maxillary anterior periapical radiographs (initial and follow-up) from 51 RET cases. Two experienced pediatric dentists established the reference standard for each fully visible tooth (n = 282) in terms of Fédération Dentaire Internationale (FDI) tooth number, RET presence/absence, apex status, periapical lesion presence/absence, and root development type (I–VI). GPT-5 and Gemini 2.5 were queried in separate reset sessions using the same standardized prompt and fixed output template. Diagnostic performance and agreement with the reference standard were assessed using standard classification metrics and agreement coefficients. Results: Both models classified FDI tooth numbers with a high degree of accuracy; Gemini 2.5 showed almost perfect agreement with the reference standard (Cohen’s κ = 0.949, observed agreement 96.5%) and outperformed GPT-5 (κ = 0.733) in terms of FDI numbering. With regard to RET detection, Gemini 2.5 achieved substantial agreement (κ = 0.796, F1 = 0.898) compared with moderate agreement in the case of GPT-5 (κ = 0.537, F1 = 0.768). In contrast, performance with regard to apex status (Fleiss’ κ = 0.270) and periapical lesion detection was limited for both models, and root development type classification was poor, with macro-averaged F1 scores of 0.130 (GPT-5) and 0.073 (Gemini 2.5) and agreement near chance. Conclusions: Off-the-shelf multimodal large language models (LLMs) demonstrated potential as assistive tools for structured, binary radiographic tasks in RET follow-up (FDI numbering and RET detection), but were not reliable in the case of complex or ordinal endpoints (apex status, lesions, root development type).


Keywords

Diagnostic accuracy; Immature permanent teeth; Large language models; Periapical radiography; Regenerative endodontic treatment


Cite and Share

Enes Mustafa Aşar,Murat Selim Botsali. Can AI see what clinicians see? A comparative study of GPT-5 and Gemini 2.5 in radiographic evaluation of regenerative endodontic treatments in immature permanent teeth. Journal of Clinical Pediatric Dentistry. 2026. 50(5);53-70.

References

[1] Umer F, Batool I, Naved N. Innovation and application of large language models (LLMs) in dentistry—a scoping review. BDJ Open. 2024; 10: 90.

[2] Çeki̇ç EC, Tavşan O. Evaluating large language models using national endodontic specialty examination questions: are they ready for real-world dentistry? BMC Medical Education. 2025; 25: 1308.

[3] Aşar EM, İpek İ, Bi Lge K. Customized GPT-4V(ision) for radiographic diagnosis: can large language model detect supernumerary teeth? BMC Oral Health. 2025; 25: 756.

[4] Salmanpour F, Akpınar M. Performance of Chat Generative Pretrained Transformer-4.0 in determining labiolingual localization of maxillary impacted canine and presence of resorption in incisors through panoramic radiographs: a retrospective study. American Journal of Orthodontics and Dentofacial Orthopedics. 2025; 168: 220–231.

[5] Silva TP, Andrade-Bortoletto MFS, Ocampo TSC, Alencar-Palha C, Bornstein MM, Oliveira-Santos C, et al. Performance of a commercially available Generative Pre-trained Transformer (GPT) in describing radiolucent lesions in panoramic radiographs and establishing differential diagnoses. Clinical Oral Investigations. 2024; 28: 204.

[6] Wei X, Yang M, Yue L, Huang D, Zhou X, Wang X, et al. Expert consensus on regenerative endodontic procedures. International Journal of Oral Science. 2022; 14: 55.

[7] Meschi N, Palma PJ, Cabanillas-Balsera D. Effectiveness of revitalization in treating apical periodontitis: a systematic review and meta-analysis. International Endodontic Journal. 2023; 56: 510–532.

[8] Erdogan O, Casey SM, Bahammam A, Son M, Mora M, Park G, et al. Radiographic evaluation of regenerative endodontic procedures and apexification treatments with the assessment of external root resorption. Journal of Endodontics. 2024; 50: 1420–1428.e1.

[9] Alfahadi HR, Al-Nazhan S, Alkazman FH, Al-Maflehi N, Al-Nazhan N. Clinical and radiographic outcomes of regenerative endodontic treatment performed by endodontic postgraduate students: a retrospective study. Restorative Dentistry & Endodontics. 2022; 47: e24.

[10] Flake NM, Gibbs JL, Diogenes A, Hargreaves KM, Khan AA. A standardized novel method to measure radiographic root changes after endodontic therapy in immature teeth. Journal of Endodontics. 2014; 40: 46–50.

[11] Buderer NM. Statistical methodology: I. Incorporating the prevalence of disease into the sample size calculation for sensitivity and specificity. Academic Emergency Medicine. 1996; 3: 895–900.

[12] AAE clinical considerations for a regenerative procedure revised 5/18/2021. 2021. Available at: https://www.aae.org/specialty/wp-content/uploads/sites/2/2021/08/ClinicalConsiderationsApprovedByREC062921.pdf (Accessed: 01 October 2025).

[13] Jiang X, Liu H, Peng C. Continued root development of immature permanent teeth after regenerative endodontics with or without a collagen membrane: a randomized, controlled clinical trial. International Journal of Paediatric Dentistry. 2022; 32: 284–293.

[14] Bani-Hani T, Wedyan M, Al-Fodeh R, Shuqeir R, Al Jundi S, Tewari N. Artificial intelligence model for application in dental traumatology. European Archives of Paediatric Dentistry. 2025; 26: 1117–1124.

[15] Taleb A, Rohrer C, Bergner B, De Leon G, Rodrigues JA, Schwendicke F, et al. Self-supervised learning methods for label-efficient dental caries classification. Diagnostics. 2022; 12: 1237.

[16] Ilhan B, Guneri P, Wilder-Smith P. The contribution of artificial intelligence to reducing the diagnostic delay in oral cancer. Oral Oncology. 2021; 116: 105254.

[17] Bilgir E, Bayrakdar İŞ, Çelik Ö, Orhan K, Akkoca F, Sağlam H, et al. An artificial intelligence approach to automatic tooth detection and numbering in panoramic radiographs. BMC Medical Imaging. 2021; 21: 124.

[18] Albitar L, Zhao T, Huang C, Mahdian M. Artificial Intelligence (AI) for detection and localization of unobturated second mesial buccal (MB2) canals in cone-beam computed tomography (CBCT). Diagnostics. 2022; 12: 3214.

[19] Qu Y, Lin Z, Yang Z, Lin H, Huang X, Gu L. Machine learning models for prognosis prediction in endodontic microsurgery. Journal of Dentistry. 2022; 118: 103947.

[20] Sadr S, Mohammad-Rahimi H, Motamedian SR, Zahedrozegar S, Motie P, Vinayahalingam S, et al. Deep learning for detection of periapical radiolucent lesions: a systematic review and meta-analysis of diagnostic test accuracy. Journal of Endodontics. 2023; 49: 248–261.e3.

[21] Patil SR, Karobari MI. Exploring artificial intelligence for enhanced endodontic practice: applications, challenges, and future directions. Advances in Public Health. 2024; 2024: 8075515.

[22] Setzer FC, Li J, Khan AA. The use of artificial intelligence in endodontics. Journal of Dental Research. 2024; 103: 853–862.

[23] Liu Z, Ai QYH, Yeung AWK, Tanaka R, Nalley A, Hung KF. Performance of a vision-language model in detecting common dental conditions on panoramic radiographs using different tooth numbering systems. Diagnostics. 2025; 15: 2315.

[24] Dursun D, Bilici Geçer R. Dental age estimation from panoramic radiographs: a comparison of orthodontist and ChatGPT-4 evaluations using the London Atlas, Nolla, and Haavikko Methods. Diagnostics. 2025; 15: 2389.

[25] Erkal D, Felek T, Butean OP, Er K. Dens invaginatus as a diagnostic challenge: evaluating large language models against expert endodontic reasoning. BMC Oral Health. 2025; 25: 1552.

[26] Lu J, Cai Q, Chen K, Kahler B, Yao J, Zhang Y, et al. Machine learning models for prognosis prediction in regenerative endodontic procedures. BMC Oral Health. 2025; 25: 234.

[27] Baraka M, El-Kateb N, Gamal M, Elwan AH, Sharaf P, Cevidanes L, et al. Assessment of volumetric changes after regenerative endodontic procedures using semiautomated and 3D U-NET automated CBCT segmentation: a retrospective cohort study. BMC Oral Health. 2025; 25: 1409.

[28] Ekmekci E, Durmazpinar PM. Evaluation of different artificial intelligence applications in responding to regenerative endodontic procedures. BMC Oral Health. 2025; 25: 53.

[29] Camlet A, Kusiak A, Ossowska A, Świetlik D. Advances in periodontal diagnostics: application of multimodal language models in visual interpretation of panoramic radiographs. Diagnostics. 2025; 15: 1851.

[30] Mine Y, Iwamoto Y, Okazaki S, Nishimura T, Tabata E, Takeda S, et al. Challenges and limitations of multimodal large language models in interpreting pediatric panoramic radiographs. International Journal of Paediatric Dentistry. 2026; 36: 74–80.

[31] AlFarabi Ali S, AlDehlawi H, Jazzar A, Ashi H, Esam Abuzinadah N, AlOtaibi M, et al. The diagnostic performance of large language models and oral medicine consultants for identifying oral lesions in text-based clinical scenarios: prospective comparative study. JMIR AI. 2025; 4: e70566.

[32] Freire Y, Santamaría Laorden A, Orejas Pérez J, Ortiz Collado I, Gómez Sánchez M, Thuissard Vasallo IJ, et al. Evaluating the influence of prompt formulation on the reliability and repeatability of ChatGPT in implant-supported prostheses. PLOS ONE. 2025; 20: e0323086.

[33] Rebitschek FG, Carella A, Kohlrausch-Pazin S, Zitzmann M, Steckelberg A, Wilhelm C. Evaluating evidence-based health information from generative AI using a cross-sectional study with laypeople seeking screening information. npj Digital Medicine. 2025; 8: 343.


JOCPD Volume 50 Issue 5 cover
Current Issue

Vol.50, Issue 5, 03 September 2026

Table of contents
All Issues

Submission Turnaround Time

Top