This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.
This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard. Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment. Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently. Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard. The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria. The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.
Study Type
OBSERVATIONAL
Enrollment
342
Large language model used to analyze standardized clinical information and periapical radiographs to provide pulpal and periapical diagnosis and endodontic case difficulty assessment according to AAE criteria
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosis
Diagnostic performance of each large language model compared with the expert consensus reference standard. Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria
Time frame: At baseline
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessment
Agreement between each large language model and the expert consensus in classifying endodontic case difficulty according to the AAE Endodontic Case Difficulty Assessment Guidelines. Overall accuracy and weighted Cohen's kappa will be calculated.
Time frame: At baseline
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.