Following model development and locking, the fixed model is evaluated in prospectively collected CT cohorts from two centers. The study is observational and does not affect clinical care. A subset of cases is used in a randomized crossover reader study.
After model locking, CT data are prospectively collected at two centers for observational validation without retraining or parameter adjustment. Model outputs do not influence patient management. A subset of eligible cases is selected for the randomized crossover reader study.
Study Type
OBSERVATIONAL
Enrollment
310
Radiologists interpret the medical images independently without any assistance from the AI model to establish a baseline performance.
Radiologists interpret the same set of medical images with the assistance of the multimodal medical imaging large model to evaluate the improvement in diagnostic performance.
The Third Affiliated Hospital of Southern Medical University
Guangzhou, Guangdong, China
Case-level Diagnostic Accuracy and Area Under the ROC Curve (AUC)
Evaluation of case-level diagnostic accuracy (defined as the proportion of diagnostic decisions matching the clinical ground-truth label) and discrimination performance (measured by AUC) to compare unaided radiologist performance versus AI-assisted performance.
Time frame: Up to 1 week per evaluation period
Diagnostic Efficiency (Reading and Reporting Time)
Measurement of diagnostic efficiency recorded as the time (in seconds) taken by radiologists to complete the case review and generate findings, with and without AI assistance.
Time frame: Up to 1 week per evaluation period
Inter-rater Agreement (Fleiss' Kappa)
Assessment of diagnostic consensus and inter-rater consistency among participating radiologists measured using Fleiss' kappa (κ).
Time frame: Up to 1 week per evaluation period
Clinical Report Quality and Semantic Accuracy Score
Assessment of AI-generated draft report quality evaluated by senior experts on a 5-point Likert scale (focusing on semantic accuracy and clinical relevance) and automated metrics (GREEN and ROUGE-L).
Time frame: Up to 1 week per evaluation period
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.