ASN's Mission

To create a world without kidney diseases, the ASN Alliance for Kidney Health elevates care by educating and informing, driving breakthroughs and innovation, and advocating for policies that create transformative changes in kidney medicine throughout the world.

learn more

Contact ASN

1401 H St, NW, Ste 900, Washington, DC 20005

email@asn-online.org

202-640-4660

The Latest on X

Kidney Week

Abstract: SA-PO0865

Artificial Intelligence (AI) Chatbots and Nephrologists' Reasoning Across Standardized Nephrology Case Vignettes

Session Information

Category: Glomerular Diseases

  • 1402 Glomerular Diseases: Clinical, Outcomes, and Therapeutics

Authors

  • Petreski, Tadej, University Medical Centre Maribor, Department of Nephrology, Maribor, Slovenia
  • Jakopin, Eva, University Medical Centre Maribor, Department of Nephrology, Maribor, Slovenia
  • Piko, Nejc, University Medical Centre Maribor, Department of Dialysis, Maribor, Slovenia
  • Vreca, Nino, University Medical Centre Maribor, Department of Nephrology, Maribor, Slovenia
  • Varda, Luka, University Medical Centre Maribor, Department of Dialysis, Maribor, Slovenia
  • Vodošek Hojs, Nina, University Medical Centre Maribor, Department of Nephrology, Maribor, Slovenia
  • Knehtl, Masa, University Medical Centre Maribor, Department of Nephrology, Maribor, Slovenia
  • Ekart, Robert, University Medical Centre Maribor, Department of Dialysis, Maribor, Slovenia
  • Bevc, Sebastjan, University Medical Centre Maribor, Department of Nephrology, Maribor, Slovenia
Background

Large language models (LLMs) are being explored for clinical decision support, but their alignment with nephrologists' reasoning in complex cases is uncertain. This study compared various LLMs to clinicians of different experience levels using standardized real-world nephrology cases, aiming to analyze agreement patterns and identify characteristics of the models' responses.

Methods

Three physician respondents—a nephrology resident, an early-career nephrologist, and a senior nephrologist—along with six LLMs (ChatGPT, Gemini, Claude, DeepSeek, Grok, and Perplexity), answered the same open-ended questions about five nephrology cases: IgA nephropathy, focal segmental glomerulosclerosis (FSGS), minimal change disease (MCD), membranous nephropathy, and ANCA vasculitis. A researcher created the case vignettes from real-world data and analyzed the responses.

Results

Output verbosity differed significantly across the models, with an average of 2251 words used for 35 answers (range from 476 by Perplexity to 5633 by Grok). Physicians used an average of 241 words. All respondents (physicians and LLMs) correctly predicted 4 of 5 kidney disease diagnoses, with everyone missing MCD in an elderly adult. However, they have included it in the differential diagnoses. Agreement on specific therapy was very high between physicians and LLMs, with LLMs placing greater emphasis on supportive therapies. For prognosis, the physician respondents have naturally predicted events using three options: "low, medium, and high", whereas the LLMs have reported a percentage range. The most discordant answers were for the probability of remission in the FSGS case, where Perplexity predicted only 10-20% compared to Gemini prediction of 70-80%, and for the probability of RRT initiation in the ANCA case, where ChatGPT predicted 30-40% compared to Claude prediction of 70-85%. There were no consistent trends in predicting more or less favorable results across LLMs.

Conclusion

The agreement between LLMs and nephrologists varied by model and question type, particularly regarding output verbosity and future event predictions. However, diagnoses, differential diagnoses, and treatment recommendations were consistently accurate and aligned with nephrologists' assessments. Future evaluations should use structured scoring rubrics to assess the clinical appropriateness, safety, and reproducibility of LLM-generated recommendations in nephrology.