RT - Journal of Clinical Pediatric Dentistry ID - 10.22514/jocpd.2026.082 T1 - Chatbots in clinical decision-making for dental trauma: from diagnosis to references A1 - Derya Sarıoğlu A1 - Büşra Yücetürk A1 - Zehra Güner A1 - Zübeyde Uçar Gündoğar K1 - Artificial intelligence; Chatbot; Dental trauma; Dentistry; Large language models; Pediatric dentistry YR - 2026 SP - 274 AB -

Background: Traumatic dental injuries are a significant public health issue, yet intelligent chatbots offering potential solutions may broaden both patients and clinicians’ horizons. This study aimed to evaluate the accuracy, temporal reliability, and reference reliability of the diagnosis and treatment responses provided by artificial intelligence (AI) chatbots to created dental trauma. Methods: 45 dental trauma scenarios based on the International Association of Dental Traumatology (IADT) guidelines were presented to four generative large language models (LLMs): ChatGPT-4o, Claude Sonnet 3.7, Gemini Advance, and DeepSeek R1. Three questions (diagnosis, treatment, and references) were asked for each scenario, and the responses were scored using a modified Global Quality Scale (mGQS) developed by the authors, allowing quantitative comparison of the results. The obtained data were analyzed using IBM SPSS Statistics version 27.0. Differences between diagnosis and treatment scores were evaluated using the Analysis of Variance (ANOVA) test, and the temporal reliability of chatbots was evaluated using the intraclass correlation coefficient (ICC). The sources provided by the LLMs were cross-checked via Google Scholar and PubMed and classified as real or fake references. Results: No significant differences were found among LLMs in diagnosis and treatment scores (p > 0.05). In the overall evaluation, DeepSeek R1 received the highest scores, while Claude Sonnet 3.7 showed the lowest average scores. When temporal reliability was assessed, ChatGPT-4o demonstrated good, clinically acceptable temporal reliability (ICC = 0.80). In contrast, Claude Sonnet showed poor reliability, while Gemini Advance and DeepSeek R1 exhibited moderate reliability. In terms of reference reliability, the highest true source rate was observed in DeepSeek R1 (84.78%), while the lowest rate was seen in Claude Sonnet 3.7 (52.57%). Conclusions: Although LLMs provided partially accurate and consistent responses for simple dental trauma cases, they are not yet suitable for clinical use, particularly in terms of treatment recommendations and source reliability.