Urological diseases such as urinary stones, prostate cancer, and bladder cancer are very common and often require highly specialized diagnosis and treatment. Today, the quality of care can vary between doctors, and there are not enough urology specialists to meet patient demand. Artificial intelligence (AI) may help doctors make faster and more consistent decisions. This study aims to develop and test an AI-powered assistant called "UroAgent" that supports doctors in diagnosing and treating urological diseases. UroAgent is built on a large language model trained specifically for urology and is connected to tools that help it retrieve medical knowledge and analyze images. To build and test UroAgent, the research team will use 1,500 past patient records from 2010-2025 and collect 500 new patient cases for validation, for a total of 2,000 cases. This is an observational study: no patient's medical treatment will be changed because of it. The goal is to create a reliable AI tool that helps improve urological care for patients.
This study protocol describes an observational study aiming to develop and validate UroAgent, an artificial-intelligence agent for the diagnosis and treatment of urological diseases. A total of 2,000 urological disease cases will be collected, comprising 1,500 retrospective cases recorded at the center between 2010 and 2025 for model development and 500 prospectively enrolled cases for independent performance validation. The primary evaluation is the concordance between UroAgent's diagnostic and treatment recommendations and the reference standards established by senior urologists, assessed through diagnostic accuracy, recommendation appropriateness, completeness, and safety; secondary evaluations include the agent's performance across disease subtypes (urinary stones, prostate cancer, bladder cancer) and its image-interpretation capability. All records will undergo de-identification, and the study will adhere to rigorous ethical standards and a pre-specified statistical analysis plan to provide robust evidence for the clinical application of this urology-specific AI agent.
Study Type
OBSERVATIONAL
Enrollment
2,000
Sun Yat-sen Memorial Hospital, Sun Yat-sen University
Guangzhou, Guangdong, China
RECRUITINGShenshan Medical Center, Sun Yat-sen Memorial Hospital, Sun Yat-sen University
Shantou, Guangdong, China
RECRUITINGGanzhou People's Hospital
Ganzhou, Jiangxi, China
RECRUITINGDiagnostic Accuracy of UroAgent
The primary outcome is UroAgent's diagnostic accuracy, measured as the F1 score of its leading diagnosis against the reference-standard final diagnosis. The reference standard is established by senior urologists from pathology, imaging, and clinical course. F1 = 2 × Precision × Recall / (Precision + Recall), computed per case and aggregated as macro-F1 across the 2,000-case cohort (1,500 retrospective + 500 prospective). Unit of measure: F1 score (range 0-1).
Time frame: Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up.
Expert Subjective Accuracy Rating
Independent urologists rate the accuracy of UroAgent's diagnosis on a 5-point Likert scale (1 = completely inaccurate, 5 = completely accurate), blinded to model identity. Reported as the mean rating and the percentage of cases rated ≥ 4; disagreements resolved by a third urologist. Unit of measure: mean rating (1-5) and percentage of cases rated accurate (%).
Time frame: Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up.
Treatment Recommendation Appropriateness of UroAgent
The proportion of cases in which UroAgent's management recommendation is rated appropriate. Two independent urologists rate each recommendation on a 5-point Likert appropriateness scale (1 = completely inappropriate, 5 = completely appropriate), blinded; "appropriate" defined as Likert ≥ 4; disagreement resolved by a third urologist. Unit of measure: percentage of cases (%) rated appropriate.
Time frame: Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up.
Clinical Safety of UroAgent Recommendations
The rate of clinically unsafe recommendations. Each case is reviewed by senior urologists using a safety rubric and flagged (binary per case) for any contraindicated, erroneous, or potentially harmful recommendation. Unit of measure: percentage of cases (%) with ≥ 1 unsafe recommendation.
Time frame: Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up.
Concordance and Non-Inferiority of UroAgent versus Clinician Diagnoses
On the same cases, UroAgent's diagnoses are compared with those of practicing urologists. Reported as the agreement rate (%) between UroAgent and clinician diagnoses, plus the diagnostic-accuracy difference in percentage points (pp); Cohen's κ reported as a supplementary statistic. Both are assessed against the reference-standard final diagnosis. Unit of measure: agreement rate (%) and accuracy difference (pp).
Time frame: Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up.
This platform is for informational purposes only and does not constitute medical advice. Always consult a qualified healthcare professional.