RT - Journal of Clinical Pediatric Dentistry ID - 10.22514/jocpd.2026.115 T1 - Can AI see what clinicians see? A comparative study of GPT-5 and Gemini 2.5 in radiographic evaluation of regenerative endodontic treatments in immature permanent teeth A1 - Enes Mustafa Aşar A1 - Murat Selim Botsali K1 - Diagnostic accuracy; Immature permanent teeth; Large language models; Periapical radiography; Regenerative endodontic treatment YR - 2026 SP - 53 AB -

Background: This study evaluates the diagnostic performance and agreement of GPT-5 and Gemini 2.5 in interpreting regenerative endodontic treatments (RET) on periapical radiographs compared with an expert reference standard. Methods: This retrospective diagnostic accuracy study included 51 paired maxillary anterior periapical radiographs (initial and follow-up) from 51 RET cases. Two experienced pediatric dentists established the reference standard for each fully visible tooth (n = 282) in terms of Fédération Dentaire Internationale (FDI) tooth number, RET presence/absence, apex status, periapical lesion presence/absence, and root development type (I–VI). GPT-5 and Gemini 2.5 were queried in separate reset sessions using the same standardized prompt and fixed output template. Diagnostic performance and agreement with the reference standard were assessed using standard classification metrics and agreement coefficients. Results: Both models classified FDI tooth numbers with a high degree of accuracy; Gemini 2.5 showed almost perfect agreement with the reference standard (Cohen’s κ = 0.949, observed agreement 96.5%) and outperformed GPT-5 (κ = 0.733) in terms of FDI numbering. With regard to RET detection, Gemini 2.5 achieved substantial agreement (κ = 0.796, F1 = 0.898) compared with moderate agreement in the case of GPT-5 (κ = 0.537, F1 = 0.768). In contrast, performance with regard to apex status (Fleiss’ κ = 0.270) and periapical lesion detection was limited for both models, and root development type classification was poor, with macro-averaged F1 scores of 0.130 (GPT-5) and 0.073 (Gemini 2.5) and agreement near chance. Conclusions: Off-the-shelf multimodal large language models (LLMs) demonstrated potential as assistive tools for structured, binary radiographic tasks in RET follow-up (FDI numbering and RET detection), but were not reliable in the case of complex or ordinal endpoints (apex status, lesions, root development type).