This prospective randomized two-period crossover reader study will evaluate whether assistance from a frozen artificial intelligence model improves physicians' interpretation of the anaerobic threshold (AT) in cardiopulmonary exercise testing (CPET). Sixteen physicians, including 8 senior and 8 junior readers, will each interpret 150 de-identified CPET examinations once without AI assistance and once with AI assistance. The primary outcome is the within-physician difference in absolute AT localization error between the assisted and unassisted conditions. Secondary outcomes include signed timing error, agreement on whether AT is present, estimates within 15 and 30 seconds of the reference, inter-reader reliability, interpretation time, and whether the effect of AI assistance differs by physician seniority.
The 150 de-identified CPET examinations will be divided into two non-overlapping case libraries, X and Y, with 75 examinations in each library. The libraries will be balanced, as far as feasible, by clinical severity and acquisition characteristics. None of the reader-study examinations will be used for model training, hyperparameter selection, or checkpoint selection. Within each seniority stratum, physicians will be randomized in a 1:1 ratio to Sequence A or Sequence B. In period 1, Sequence A will interpret library X without AI assistance and library Y with AI assistance; Sequence B will receive the opposite conditions. Following a 14 ± 3 day washout, the assistance conditions will be reversed in period 2. Case order will be randomized within each period, and new blinded task identifiers will be assigned in period 2. Thus, every physician will interpret every examination once without AI assistance and once with AI assistance, yielding 4,800 interpretations. In the assisted condition, the interface will display the frozen model's predicted AT time. In the unassisted condition, no model output will be shown. Readers will be masked to case source, routine-report and reference-standard labels, other readers' annotations, and their own annotations from the previous period. For each examination, the reader will classify AT as confirmed, not reached, or unable to confirm. When AT is confirmed, the reader will record the AT time and the interpretive method used. Interpretation time will also be recorded. The reference standard will be prespecified from unassisted senior-physician annotations. AT will be considered present when at least 6 of 8 senior physicians confirm it, and the reference AT time will be the median of the confirming senior estimates. AT will be considered not reached when at least 6 of 8 senior physicians make that classification. Discordant cases will undergo adjudication by an additional senior physician who is masked to model output and reader identity. For analyses of senior readers, a leave-one-reader-out reference will be used. The primary analysis will compare absolute AT localization error between assisted and unassisted interpretations using a crossed mixed-effects model. Fixed effects will include assistance condition, physician seniority, the assistance-by-seniority interaction, period, randomized sequence, and case library. Random intercepts will be included for physician and examination. Estimates will be reported with 95% confidence intervals. No simulated or pilot-result data will be included in the prospective study analysis.
Study Type
INTERVENTIONAL
Allocation
RANDOMIZED
Purpose
DIAGNOSTIC
Masking
NONE
Enrollment
16
A frozen PACE-Former model provides a predicted anaerobic threshold time within the CPET review interface. Physicians retain responsibility for the final classification and AT time annotation. No other model output or reference label is displayed.
The physician reviews the same de-identified CPET data using the study interface without display of any model prediction or AI-generated suggestion and independently classifies AT status and records the AT time when confirmed.
Zhongshan Hospital, Fudan University
Shanghai, Shanghai Municipality, China
Difference in Absolute Anaerobic Threshold Localization Error
Absolute difference, in seconds, between the physician-selected AT time and the prespecified expert reference AT time. The primary estimand is the within-physician difference between AI-assisted and unassisted interpretations, analyzed across the same examinations using a crossed mixed-effects model.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Signed Anaerobic Threshold Localization Error
Signed difference, in seconds, between the physician-selected AT time and the prespecified expert reference AT time. Negative values indicate earlier placement and positive values indicate later placement relative to the reference.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Proportion of Anaerobic Threshold Estimates Within 15 Seconds of the Reference
Proportion of interpretable examinations for which the physician-selected AT time is within 15 seconds, in either direction, of the prespecified expert reference AT time.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Proportion of Anaerobic Threshold Estimates Within 30 Seconds of the Reference
Proportion of interpretable examinations for which the physician-selected AT time is within 30 seconds, in either direction, of the prespecified expert reference AT time.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Agreement With the Reference Anaerobic Threshold Status
Agreement between the physician classification (AT confirmed, AT not reached, or unable to confirm) and the prespecified expert reference classification. Results will be summarized as overall agreement and condition-specific classification performance.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Inter-Reader Reliability
Reliability among physicians for AT timing and AT-status classification under AI-assisted and unassisted conditions, summarized using appropriate intraclass correlation and chance-corrected agreement statistics with 95% confidence intervals.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Interpretation Time per Examination
Elapsed time, in seconds, from opening the blinded CPET task to submission of the physician's AT classification and, when applicable, the selected AT time.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
Interaction Between AI Assistance and Physician Seniority
Difference in the effect of AI assistance between senior and junior physicians, estimated from the prespecified assistance-by-seniority interaction term in the crossed mixed-effects model for absolute AT localization error.
Time frame: Recorded immediately after each case interpretation and analyzed after completion of both study periods, up to 2 months
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.