The goal of this clinical trial is to learn whether kidney transplant follow-up visits run by a health professional using an artificial intelligence (AI) computer helper are as good as usual follow-up visits with a doctor. People who receive a kidney transplant need check-up visits for the rest of their lives. These visits help keep the new kidney working well. As more people live longer with a transplant, clinics get busier. An AI helper may make visits more consistent and free up doctor time. This idea has not yet been tested well in real clinics. The main questions this study will answer are: Are AI-assisted visits as good as usual doctor visits? Trained reviewers will rate the quality of each visit. The reviewers will not know which type of visit they are rating. Are patients as satisfied with their visit? How long does the visit take, and how much time does the clinician spend on paperwork? Do the visits do a better job of checking key items, such as screening tests, vaccines, and transplant medicines? Are the visits safe for patients? Adults can take part if they have a working kidney transplant, if the transplant was at least 6 months ago, and if they can complete a routine visit on their own. Researchers will place 140 adults into two groups by chance, like flipping a coin. One group will have a visit run by a health professional who uses the AI helper. The other group will have a usual visit with a doctor. In both groups, the AI cannot order tests, change medicines, or make a diagnosis on its own. A clinician checks and approves everything. Participants will take part in one study follow-up visit. The visit will be audio-recorded so reviewers can rate it later. After the visit, participants will fill out a short satisfaction survey on a tablet.
Rationale. Kidney transplant recipients require lifelong, protocol-driven outpatient follow-up. After the first 6 to 12 months, care of stable recipients becomes standardized, and a large share of each routine visit is spent on standardizable tasks: laboratory review, immunosuppression reconciliation, screening and vaccination checks, and clinical documentation. Growing recipient numbers, a shortage of nephrologists, and the concentration of specialists in tertiary centers lengthen follow-up intervals, particularly within the Brazilian Unified Health System (SUS). Large language models (LLMs) have shown promise on patient-facing tasks, but existing evidence derives mainly from vignette- or forum-based comparisons, without real-workflow evaluation, prospective safety endpoints, or blinded quality instruments. A retrieval-augmented generation (RAG) architecture that grounds model outputs in an indexed base of guidelines and institutional protocols, with verifiable citations and explicit uncertainty flagging, addresses hallucination risk and enables auditable, specialty-specific use. This trial moves the evaluation of LLM support from vignettes to a real ambulatory workflow under mandatory human supervision. Study design and setting. This is a randomized, controlled, parallel-group, non-inferiority trial conducted at the kidney transplant outpatient clinic of the Faculdade de Medicina de Botucatu, Universidade Estadual Paulista (UNESP), Brazil. Reporting follows the CONSORT 2010 statement. Allocation is 1:1 to an AI-assisted follow-up consultation or a usual physician-led consultation. Intervention (AI-assisted arm). A trained health professional conducts the routine follow-up consultation while a specialized LLM assistant provides real-time support. The assistant is built on a general-purpose LLM adapted for clinical dialogue and operates through a RAG pipeline that indexes KDIGO guidelines and clinic standard operating procedures. During the encounter the assistant supplies a structured content checklist, guideline-anchored suggestions with source citations, uncertainty flags, and a draft structured clinical note and orders. Guardrails prevent the assistant from finalizing medication changes, orders, or diagnoses. Every proposed action is flagged for clinician review; the responsible clinician verifies vital signs and findings, corrects inaccuracies, and co-signs the encounter before any action is enacted (human-in-the-loop). Transcripts and structured data are stored in the electronic data capture system. Comparator (usual-care arm). Standard physician-led transplant follow-up per routine clinic practice and current guidelines, comprising history, medication reconciliation with a focus on immunosuppression, focused examination, and management plan. To limit performance bias, both arms follow an identical pre-specified content checklist and target a similar consultation duration. All consultations are audio-recorded to permit later blinded assessment. Randomization, allocation concealment, and masking. A computer-generated randomization sequence with variable block sizes is implemented in REDCap; sequentially numbered, opaque, sealed envelopes serve as an offline contingency. Allocation remains concealed until assignment. Because the nature of the intervention precludes masking of participants and treating clinicians, masking is applied to the outcome assessors and to the statistician. Trained raters who score consultation quality are blinded to allocation, as is the analyst. Sample size. The trial is powered on the primary outcome, the MAAS-Global total score (range 0 to 6). Using a pre-specified non-inferiority margin of 0.5 points, an assumed common standard deviation of 1.0, one-sided alpha of 0.025, 80% power, and 1:1 allocation, 63 participants per arm are required. Allowing for 10% loss or withdrawal yields 70 participants per arm, for a planned total of 140. Statistical analysis. Analyses follow the intention-to-treat principle. Continuous outcomes are analyzed with linear regression or mixed-effects models, with a fixed effect for group, pre-specified optional covariates (age, sex, and time since transplant), and, where applicable, a random intercept for rater. Non-inferiority for the primary outcome is assessed from the two-sided 95% confidence interval for the adjusted mean difference (AI-assisted minus usual care); non-inferiority is concluded if the lower confidence bound exceeds the margin of -0.5. Binary outcomes are compared using risk differences with 95% confidence intervals. Missing data are addressed with multiple imputation when missingness exceeds 5% and is plausibly missing at random. Two-sided p-values are reported, with emphasis on effect sizes and confidence intervals. A detailed statistical analysis plan is finalized before database lock. Data management and quality assurance. Study data are collected and managed in REDCap. To estimate inter-rater reliability, two independent raters score 20% of consultations in duplicate; disagreement greater than 1 point on any item triggers adjudication and rater retraining. The patient satisfaction questionnaire is self-administered on a tablet immediately after the consultation, without staff assistance, to preserve independence. Consultation timing data are captured for each encounter to support the efficiency analyses. Data handling complies with the Brazilian General Data Protection Law (LGPD), and safety is monitored continuously throughout enrollment. Duration. The project is planned for 36 months, spanning system development and RAG indexing, cross-cultural adaptation of the satisfaction instrument, rater training and calibration, a stabilization pilot, enrollment and data collection, analysis, and dissemination.
Study Type
INTERVENTIONAL
Allocation
RANDOMIZED
Purpose
HEALTH_SERVICES_RESEARCH
Masking
SINGLE
Enrollment
140
A specialized large language model (LLM) assistant supports the consulting health professional in real time during a routine kidney transplant follow-up consultation. The assistant uses a retrieval-augmented generation (RAG) pipeline indexed on KDIGO guidelines and clinic standard operating procedures, and provides a structured content checklist, guideline-anchored suggestions with source citations, uncertainty flags, and a draft structured clinical note and orders. Guardrails prevent the assistant from finalizing medication changes, orders, or diagnoses. The responsible clinician reviews and verifies all findings, corrects inaccuracies, and approves and co-signs the encounter before any action is enacted (human-in-the-loop).
Standard physician-led kidney transplant follow-up according to routine clinic practice and current guidelines, without the AI assistant. The consultation comprises history-taking, medication reconciliation with a focus on immunosuppression, focused examination, and a management plan, following the same pre-specified content checklist as the experimental arm.
UPECLIN
Botucatu, São Paulo, Brazil
Consultation quality measured by the MAAS-Global global score
Overall quality of the follow-up consultation, scored with the MAAS-Global instrument, a validated tool for rating clinician consultation skills. Consultation quality is reported as the MAAS-Global global score, which ranges from 0 to 6, where 0 indicates the poorest consultation quality and 6 indicates the highest consultation quality (higher scores indicate better quality). Consultations are audio-recorded and scored by trained raters who are blinded to group allocation. This is the primary endpoint for the non-inferiority comparison between the AI-assisted consultation and the usual physician-led consultation.
Time frame: Consultation quality measured at Day 1
Patient satisfaction measured by the VSQ-9
Participant satisfaction with the consultation, measured with the Visit-Specific Satisfaction Questionnaire (VSQ-9), self-administered on a tablet immediately after the visit without staff assistance. Each item is scored on a 5-point scale and transformed to a 0 to 100 metric; the total score is the mean of the answered items and ranges from 0 to 100, where 0 indicates the lowest satisfaction and 100 indicates the highest satisfaction (higher scores indicate greater satisfaction).
Time frame: Patient satisfaction at day 1
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.