Cardiopulmonary exercise testing requires clinicians to integrate multiple physiological measurements and explain their meaning, limitations, safety implications, and appropriate next steps. This randomized two-period crossover study will evaluate whether physician-supervised GPT-assisted drafting preserves the clinical acceptability of final patient-facing responses while reducing task completion time compared with physician-only drafting. Licensed physicians or institutionally recognized clinical trainees will respond in Chinese to 20 synthetic cardiopulmonary exercise testing scenarios. Each participant will complete 10 scenarios in each condition using two non-overlapping case sets. In the GPT-assisted condition, a frozen model-generated draft will be displayed and may be edited, deleted, or completely rewritten by the physician. The primary hypothesis is that GPT-assisted drafting is noninferior to physician-only drafting for clinical acceptability, using a noninferiority margin of 5 percentage points. Completion time will be evaluated for superiority only if noninferiority is established.
Before the human-participant crossover study, four large language models were evaluated in a separate model-comparison phase using a non-overlapping 60-case benchmark. One model was selected according to prespecified criteria addressing safety, clinical acceptability in high-risk cases, and performance consistency. This model-comparison phase was completed before trial registration and is not part of the prospectively registered participant trial. Before participant enrollment, the selected model's first technically valid response to each of 20 held-out synthetic cardiopulmonary exercise testing cases will be frozen. The exact model identifier, version or snapshot, system prompt, user prompt, generation parameters, generation date, and output-integrity information will be documented in the study records. The human-participant study uses a prospective, randomized, two-period, two-treatment crossover design. Participants will be assigned to one of four sequences balancing study-condition order and case-set allocation. Each participant will complete 10 cases under physician-only drafting and 10 cases under GPT-assisted drafting. The two periods will be separated by 7 plus or minus 2 days. In the physician-only condition, participants will receive a blank response field. In the GPT-assisted condition, participants will receive a frozen GPT-generated draft that they may accept, edit, delete, or completely rewrite. The physician remains responsible for the submitted final response. Each case has a maximum completion time of 8 minutes. A reminder will be displayed after 6 minutes, and the response will be submitted automatically at 8 minutes. Participants will receive a fixed 3-minute rest after every five cases. Use of external websites, additional generative artificial intelligence tools, clinical guidelines, personal notes, or consultation with another person is prohibited during study tasks. The initial planned enrollment is 80 participants, with 20 participants allocated to each sequence. A blinded sample size re-estimation will be conducted after 40 evaluable participants have completed both periods. If required, enrollment may be increased in blocks of eight participants to a maximum of 112, subject to ethics approval and prospective registry updating. All study cases are synthetic. No real patient records will be used. Individual physician performance will not be disclosed to employers or supervisors and will not be used for employment or professional evaluation.
Study Type
INTERVENTIONAL
Allocation
RANDOMIZED
Purpose
HEALTH_SERVICES_RESEARCH
Masking
SINGLE
Enrollment
80
A frozen GPT-generated draft will be displayed for each synthetic case. The participant may accept, edit, delete, or completely rewrite the draft before submitting the final response. The participant remains responsible for the final response. Each task has a maximum duration of 8 minutes, and external information sources or additional artificial intelligence tools are prohibited.
The participant will receive a blank response field and will independently prepare the final response. Each task has a maximum duration of 8 minutes, and external information sources or artificial intelligence tools are prohibited.
Zhongshan Hospital, Fudan University
Shanghai, Shanghai Municipality, China
Proportion of Final Responses Meeting the Composite Clinical Acceptability Criterion
A task-level binary outcome. A final response is classified as clinically acceptable only when all of the following criteria are met: no S2 or S3 safety error; inclusion of all case-specific critical facts; inclusion of at least 75% of general required facts; an accuracy score of at least 4 on a 5-point scale; and a communication score of at least 3 on a 5-point scale. A higher proportion indicates better performance. The primary comparison is the marginal absolute probability difference between GPT-assisted and physician-only drafting, with a noninferiority margin of minus 5 percentage points.
Time frame: During each study period, with the second period occurring 5 to 9 days after the first period
Task Completion Time
Time in seconds from display of the case and response interface to submission of the final response. Responses automatically submitted at the 8-minute limit will be recorded as 480 seconds. A shorter completion time indicates greater efficiency.
Time frame: During each study period, with the second period occurring 5 to 9 days after the first period
Proportion of Responses Containing an S2 Safety Error
An S2 event is a clinically important error or omission with a plausible potential to alter clinical management. Each response will be classified by masked outcome assessors using the prespecified safety rubric.
Time frame: During each study period, with the second period occurring 5 to 9 days after the first period
Proportion of Responses Containing an S3 Safety Error
An S3 event is an error or omission with a plausible potential for immediate serious harm. Each response will be classified by masked outcome assessors using the prespecified safety rubric.
Time frame: During each study period, with the second period occurring 5 to 9 days after the first period
Clinical Accuracy Score
Accuracy of the final response assessed by masked outcome assessors on a 5-point scale. Higher scores indicate greater clinical accuracy.
Time frame: During each study period, with the second period occurring 5 to 9 days after the first period
Communication Quality Score
Communication quality of the final response assessed by masked outcome assessors on a 5-point scale. Higher scores indicate clearer and more appropriate patient-facing communication.
Time frame: During each study period, with the second period occurring 5 to 9 days after the first period
Raw NASA Task Load Index Score
Workload will be assessed using the six Raw NASA Task Load Index dimensions, each scored from 0 to 100. The unweighted mean of the six dimensions will be calculated. Higher scores indicate greater perceived workload.
Time frame: Immediately after each 10-case study period, with study periods occurring 5 to 9 days apart
Perceived Helpfulness of the GPT-Assisted Workflow
Participant-reported helpfulness rating on a 5-point Likert scale. Higher scores indicate greater perceived helpfulness.
Time frame: Immediately after completion of the GPT-assisted study period
Trust in the GPT-Assisted Workflow
Participant-reported trust rating on a 5-point Likert scale. Higher scores indicate greater trust.
Time frame: Immediately after completion of the GPT-assisted study period
Perceived Risk of the GPT-Assisted Workflow
Participant-reported risk rating on a 5-point Likert scale. Higher scores indicate greater perceived risk.
Time frame: Immediately after completion of the GPT-assisted study period
Willingness to Use the GPT-Assisted Workflow Under Physician Supervision
Participant-reported willingness rating on a 5-point Likert scale. Higher scores indicate greater willingness to use the workflow under physician supervision.
Time frame: Immediately after completion of the GPT-assisted study period
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.