The Movement Disorders Society (MDS) Unified Parkinson's Disease Rating Scale (UPDRS) Part III (MDS-UPDRS III) is the primary assessment method for motor symptoms in Parkinson's disease patients. Currently, movement disorder specialists conduct semi-quantitative scoring, which entails limitations such as subjectivity, weak sensitivity, and a limited number of professional physicians. This study, based on machine vision, establishes gold standard labels according to expert scoring. By using machine learning, we develop a machine rating model and compare the model's performance with gold standard rating and general clinical rating to investigate the accuracy of machine vision-based MDS-UPDRS III machine rating.
Study Type
OBSERVATIONAL
Enrollment
1,679
Patients' performance of MDS-UPDRS III will be recorded.
Center for Movement Disorders, Department of Neurology, Beijing Tiantan Hospital, Capital Medical University
Beijing, Beijing Municipality, China
Beijing Hospital, Neurology Department
Beijing, Beijing Municipality, China
Department of Neurology, Fujian Medical University Union Hospital
Fuzhou, Fujian, China
Department of Neurology, Guangdong Neuroscience Institute, Guangdong General Hospital, Guangdong Academy of Medical Sciences
Guangzhou, Guangdong, China
Department of Neurology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology
Wuhan, Hubei, China
Department of Neurology and Clinical Research Center of Neurological Disease, The Second Affiliated Hospital of Soochow University
Suzhou, Jiangsu, China
Department of Neurology and Institute of Neurology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine
Shanghai, Shanghai Municipality, China
Department of Neurology, West China Hospital, Sichuan University
Chengdu, Sichuan, China
Item-level: MAE
At the item level, mean absolute error (MAE) between paired AI scores and consensus reference scores, calculated as the average absolute difference across items. Lower values indicate closer agreement.
Time frame: 1 day
Item-level: P(|Δ|≥2)
At the item level, the proportion of paired AI and consensus reference scores with an absolute difference of 2 or more points. Lower values indicate fewer large item-level scoring disagreements.
Time frame: 1 day
Score-level: ICC
Intraclass correlation coefficient (ICC) between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. Higher ICC values indicate greater agreement.
Time frame: 1 day
Score-level: SEM
Standard error of measurement (SEM) for AI-derived scores relative to consensus reference scores, calculated separately for each of the four subdomain scores and the total score. SEM quantifies measurement error in the units of the corresponding score, with lower values indicating greater measurement precision.
Time frame: 1 day
Score-level: SDC
Smallest detectable change (SDC), derived from the measurement error and calculated separately for each of the four subdomain scores and the total score. SDC represents the minimum score change required to exceed expected measurement error. Lower values indicate greater measurement precision.
Time frame: 1 day
Item-level: Within-one agreement (ACC1; P(|Δ| ≤ 1))
At the item level, ACC1 is defined as the proportion of paired AI and consensus reference scores with an absolute difference of no more than 1 point, i.e., P(\|Δ\| ≤ 1). Higher values indicate closer item-level agreement.
Time frame: 1 day
Item-level: Exact-error proportions [P(|Δ| = k)]
At the item level, exact-error proportions, P(\|Δ\| = k), represent the proportions of paired AI and consensus reference scores with each exact absolute error magnitude k. In particular, P(\|Δ\| = 0) corresponds to exact-match accuracy (ACC), defined as the proportion of items for which the AI score exactly matches the consensus reference score. The remaining values of k characterize the distribution of item-level scoring errors.
Time frame: 1 day
Item-level: Per-class exact recall
At the item level, for each consensus reference score class, the proportion of items for which the AI score exactly matches the consensus reference score among all items belonging to that reference class. Higher values indicate better class-specific exact agreement.
Time frame: 1 day
Item-level: Confusion matrix
At the item level, a cross-tabulation of AI scores against consensus reference scores, showing the number of paired ratings for each combination of reference and AI score categories. Rows represent consensus reference scores and columns represent AI scores.
Time frame: 1 day
Item-level: Row-normalised confusion matrix
At the item level, the confusion matrix normalised within each consensus reference score row so that each row sums to 1, showing the distribution of AI scores conditional on each reference score class.
Time frame: 1 day
Score-level: Limits of agreement
Bland-Altman limits of agreement between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. The limits characterize the range within which most paired differences between AI and reference scores are expected to fall.
Time frame: 1 day
Score-level: MAE
Mean absolute error (MAE) between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score as the average absolute difference between paired scores. Lower values indicate smaller scoring errors.
Time frame: 1 day
Score-level: RMSE
Root mean square error (RMSE) between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. RMSE gives greater weight to larger scoring errors, with lower values indicating closer agreement.
Time frame: 1 day
Score-level: Spearman correlation
Spearman rank correlation between AI-derived scores and consensus reference scores, calculated separately for each of the four subdomain scores and the total score. Higher values indicate a stronger monotonic association between AI-derived and reference scores.
Time frame: 1 day
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.