Comparative Assessment of Performance of Domain-Specific and Primed Versus Non-Primed Artificial Intelligence Chatbots for Clinical Decision-Making in Injuries to the Primary Dentition: An Exploratory Pilot Study
Dental Traumatology, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Basım Tarihi: 2026
- Doi Numarası: 10.1111/edt.70110
- Dergi Adı: Dental Traumatology
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, CINAHL, EMBASE, MEDLINE, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
- Uşak Üniversitesi Adresli: Evet
Özet
Background/Aims: To compare the performance of a domain-specific dental trauma chatbot with general-purpose artificial intelligence chatbots under primed and non-primed conditions for clinical decision-making in injuries to the primary dentition. Methods: Fifteen standardized clinical case scenarios were developed for evaluating three chatbot systems: a domain-specific model (Dental Trauma Evo) and two general-purpose models (ChatGPT 5.2 and Perplexity Pro). General-purpose chatbots were evaluated under primed and non-primed conditions, where priming involved providing a summarized guideline document prior to scenario input. Chatbots were required to generate responses addressing diagnosis, immediate management, follow-up intervals, and radiographic recommendations. Two Pediatric dentists independently evaluated responses using a binary scoring system based on International Association of Dental Traumatology (IADT) guidelines. Statistical comparisons were performed using McNemar's test and agreement analysis with kappa statistics. Results: All chatbots demonstrated complete diagnostic accuracy across scenarios. The domain-specific chatbot achieved 100% accuracy in immediate management decisions, while minor inaccuracies were observed among general-purpose models. The most substantial differences were observed in follow-up interval and radiographic recommendations. Priming significantly improved the performance of general-purpose chatbots, with ChatGPT and Perplexity demonstrating marked gains in follow-up interval accuracy and radiographic recommendations. In these domains, primed general-purpose models demonstrated greater agreement with IADT guideline recommendations than the domain-specific chatbot. Conclusions: Guideline-based priming was associated with improved performance of general-purpose AI chatbots in clinical decision-making for injuries to the primary dentition, particularly for follow-up scheduling and radiographic recommendations. While domain-specific models provide reliable management guidance, contextual exposure to clinical guidelines enables general large language models to demonstrate a high level of agreement with established recommendations in certain aspects of trauma care in primary dentition. These findings highlight the potential role of prompt-guided AI systems as supportive tools in dental traumatology decision-making.