This observational and methodological study aims to compare the performance of large language models in generating electrode contact configuration recommendations for epidural electrical stimulation in spinal cord injury. Five standardized synthetic spinal cord injury scenarios will be presented to four large language models: ChatGPT-4o, Claude, Grok 3, and Gemini 2.5 Pro. Each model will receive the same standardized prompt. The generated responses will be anonymized and evaluated independently by experts with experience in spinal cord injury rehabilitation and epidural electrical stimulation. The responses will be assessed in five main areas: clinical accuracy, technical feasibility, safety awareness, consistency with current clinical guidance, and completeness of the response. Agreement between expert evaluators will also be examined. No real patients, human participants, clinical interventions, or personal health data are included in this study. The study is designed to explore the potential and current limitations of large language models as artificial intelligence-based clinical decision-support tools in neurorehabilitation.
Study Type
OBSERVATIONAL
Enrollment
20
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
The large language model receives five standardized synthetic spinal cord injury scenarios using an identical standardized prompt and generates recommendations for epidural electrical stimulation electrode contact configuration mapping. No intervention is administered to human participants.
Istanbul Gelisim University
Istanbul, Istanbul, Turkey (Türkiye)
Clinical Accuracy Score of Large Language Model Responses
Clinical accuracy of the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater clinical accuracy of the generated recommendations.
Time frame: At the time of expert evaluation, within 1 week after study initiation
Technical Feasibility Score of Large Language Model Responses
The technical feasibility of epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater technical feasibility and applicability of the generated recommendations.
Time frame: At expert evaluation, within 1 week after study initiation
Safety Awareness Score of Large Language Model Responses
The safety awareness demonstrated in the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater recognition and consideration of relevant safety issues.
Time frame: At expert evaluation, within 1 week after study initiation
Clinical Guideline Consistency Score of Large Language Model Responses
The consistency of the generated epidural electrical stimulation electrode contact configuration recommendations with current clinical guidance will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate greater consistency with current clinical guidance and relevant evidence-based recommendations.
Time frame: At expert evaluation, within 1 week after study initiation
Response Completeness Score of Large Language Model Responses
The completeness of the epidural electrical stimulation electrode contact configuration recommendations generated by each large language model will be independently evaluated by expert reviewers using a 5-point Likert-type rating scale. Higher scores indicate more complete and comprehensive responses.
Time frame: At expert evaluation, within 1 week after study initiation
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.