Large Language Models for Traumatic Dental Injuries Across Web-Based and Mobile-Based Interfaces: Assessing Accuracy, Quality, and Temporal Consistency
Clinical and Experimental Dental Research, cilt.12, sa.4, 2026 (ESCI, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 12 Sayı: 4
- Basım Tarihi: 2026
- Doi Numarası: 10.1002/cre2.70416
- Dergi Adı: Clinical and Experimental Dental Research
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, EMBASE, MEDLINE, Directory of Open Access Journals, Health Research Premium Collection (ProQuest)
- Anahtar Kelimeler: artificial intelligence, dentistry, endodontics, information reliability, natural language processing, traumatic dental injuries
- Uşak Üniversitesi Adresli: Evet
Özet
Objectives: Traumatic dental injuries (TDIs) are frequent in clinical practice and require rapid, guideline-based decisions, yet accessing accurate and reliable information may be challenging. Large language models (LLMs) such as ChatGPT, Gemini, DeepSeek, and Qwen are increasingly used as quick online information tools; however, evidence regarding their accuracy, consistency, and the influence of different user interfaces is limited. This study aimed to evaluate the performance of several LLMs in answering TDI-related questions through both web-based interfaces and mobile phone applications. Material and Methods: Twenty questions were prepared according to the 2020 International Association of Dental Traumatology (IADT) guidelines, including 10 open-ended and 10 yes–no items. Four LLMs (ChatGPT-4o, DeepSeek-V3, Gemini 2.0 Flash, Qwen2.5-Max) were queried simultaneously via web and mobile interfaces over five consecutive days, generating 800 responses. Open-ended answers were assessed using the Global Quality Score (GQS) and modified DISCERN (mDISCERN), while yes–no responses were compared with a predetermined answer key. Statistical analyses were performed using IBM SPSS v23.0, with significance set at p < 0.05. Results: Qwen2.5-Max demonstrated comparatively higher GQS and mDISCERN scores across both interfaces. Accuracy for yes–no questions ranged from 86% to 91% without significant differences among models. Interface comparisons showed that ChatGPT-4o generated comparatively higher-quality responses on the web, whereas Qwen2.5-Max performed better on mobile. Over the 5-day period, Qwen2.5-Max showed relatively higher temporal consistency, while DeepSeek-V3 exhibited notable day-to-day variation. Conclusions: LLMs may serve as useful supplementary tools for providing guideline-based information on TDIs, especially for straightforward, closed-ended clinical questions. However, their performance varies by model, interface, and question type. Qwen2.5-Max demonstrated comparatively higher performance across several evaluated measures. Despite these results, LLM-generated information should be interpreted cautiously and verified by dental professionals before being used in clinical decision-making.