A multicentre, randomised diagnostic accuracy study to evaluate whether the rare disease-specific AI can improve diagnostic accuracy and efficiency for physicians managing real-world clinical cases.
Rare diseases collectively affect approximately 300 million individuals worldwide. This prolonged diagnostic delay is attributable in large part to the breadth of over 7,000 recognized rare conditions, which far exceeds the clinical exposure of any individual physician. A rare disease-specific diagnostic AI was developed by Peking Union Medical College Hospital (PUMCH), supporting differential diagnosis generation, clinical workup planning, and genomic variant interpretation. A balanced crossover design ensures that each enrolled physician serves as their own control, substantially reducing confounding from inter-reader variability in baseline diagnostic competency. Within each physician, cases are randomly assigned at the case level to either the AI-assisted or unassisted condition, such that each physician reads a subset of cases with AI assistance and the remaining cases without. This within-reader, case-level randomization eliminates the need for a washout period and directly controls for inter-reader differences in baseline diagnostic competency. All cases are collected from real-world clinical settings with independently confirmed gold-standard diagnoses and span a pre-specified spectrum of rare and non-rare disease categories, reflecting the differential diagnostic challenge encountered in routine clinical practice, to ensure diagnostic breadth and clinical representativeness. Physician seniority (junior vs. senior) is incorporated as a pre-specified stratification and subgroup analysis variable. Diagnostic outputs are evaluated by an independent Expert Adjudication Committee, blinded to the assistance condition, using standardized scoring criteria established prior to data collection.
Study Type
INTERVENTIONAL
Allocation
RANDOMIZED
Purpose
DIAGNOSTIC
Masking
SINGLE
Enrollment
150
A rare disease-specific diagnostic AI model is used to accept free text input and assist in rare disease diagnoses. During the experimental condition, physicians may interact with the system freely alongside standard clinical resources to support their diagnostic reasoning.
Peking Union Medical College Hospital
Beijing, China
Cangzhou Central Hospital
Cangzhou, China
Changchun Sacred Heart Hospital
Changchun, China
Top-3 Diagnostic Accuracy
The percentage of definitive diagnosis is included within the physician's top 3 choices.
Time frame: Up to 60 minutes per case (from case presentation to diagnostic report submission).
Diagnosis Time per Case
Elapsed time from initial case presentation to final diagnostic report submission, recorded automatically via system logs.
Time frame: Up to 60 minutes per case (from case presentation to diagnostic report submission).
Workup Plan Quality
Quality score of the clinical workup plan assigned by an independent expert committee using a standardized Likert Scale. Scores range from 1 to 10, with higher scores indicating better workup plan quality.
Time frame: Up to 60 minutes per case (from case presentation to diagnostic report submission).
Physician Reported Usability of the AI-Assisted Diagnostic System
Physician-reported usability of the AI system, assessed after completion of each AI-assisted case reading using a 10-point physician-rated usability scale. Scores range from 1 to 10, with higher scores indicating better system usability.
Time frame: Up to 60 minutes per case (upon completion of each case reading).
Physician Reported Workload
Task-related workload experienced by physicians, assessed after completion of each AI-assisted case reading using a 10-point Physician Workload Likert scale. Scores range from 1 to 10, with higher scores indicating a higher workload.
Time frame: Up to 60 minutes per case (upon completion of each case reading).
Physician Satisfaction
Overall satisfaction of physicians with the diagnostic workflow, assessed after completion of each AI-assisted case reading using a 10-point Satisfaction Likert scale. Scores range from 1 to 10, with higher scores indicating higher satisfaction.
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.
Dongguan People's Hospital
Dongguan, China
First People's Hospital of Foshan
Foshan, China
Guizhou Provincial People's Hospital
Guiyang, China
Jilin Central General Hospital
Jilin City, China
The First People's Hospital of Yunnan Province
Kunming, China
Tibet Autonomous Region People's Hospital
Lhasa, China
Tianjin Children's Hospital
Tianjin, China
...and 3 more locations
Time frame: Up to 60 minutes per case (upon completion of each case reading).
Physician Intention to Adopt AI-Assisted Diagnostic Support
Physician willingness to integrate AI system into routine clinical practice, assessed after completion of each AI-assisted case reading using a 10-point Adoption Intention Likert scale. Scores range from 1 to 10, with higher scores indicating higher adoption intention.
Time frame: Up to 60 minutes per case (upon completion of each case reading).