Investigating the Impact of AI on Shared Decision-Making in Post-Kidney Transplant Care (PRIMA-AI): Protocol for a Randomized Controlled Trial
Bibliographic record
Abstract
Background Patients after kidney transplantation eventually face the risk of graft loss with the concomitant need for dialysis or retransplantation. Choosing the right kidney replacement therapy after graft loss is an important preference-sensitive decision for kidney transplant recipients. However, the rate of conversations about treatment options after kidney graft loss has been shown to be as low as 13% in previous studies. It is unknown whether the implementation of artificial intelligence (AI)–based risk prediction models can increase the number of conversations about treatment options after graft loss and how this might influence the associated shared decision-making (SDM). Objective This study aims to explore the impact of AI-based risk prediction for the risk of graft loss on the frequency of conversations about the treatment options after graft loss, as well as the associated SDM process. Methods This is a 2-year, prospective, randomized, 2-armed, parallel-group, single-center trial in a German kidney transplant center. All patients will receive the same routine post–kidney transplant care that usually includes follow-up visits every 3 months at the kidney transplant center. For patients in the intervention arm, physicians will be assisted by a validated and previously published AI-based risk prediction system that estimates the risk for graft loss in the next year, starting from 3 months after randomization until 24 months after randomization. The study population will consist of 122 kidney transplant recipients >12 months after transplantation, who are at least 18 years of age, are able to communicate in German, and have an estimated glomerular filtration rate <30 mL/min/1.73 m2. Patients with multi-organ transplantation, or who are not able to communicate in German, as well as underage patients, cannot participate. For the primary end point, the proportion of patients who have had a conversation about their treatment options after graft loss is compared at 12 months after randomization. Additionally, 2 different assessment tools for SDM, the CollaboRATE mean score and the Control Preference Scale, are compared between the 2 groups at 12 months and 24 months after randomization. Furthermore, recordings of patient-physician conversations, as well as semistructured interviews with patients, support persons, and physicians, are performed to support the quantitative results. Results The enrollment for the study is ongoing. The first results are expected to be submitted for publication in 2025. Conclusions This is the first study to examine the influence of AI-based risk prediction on physician-patient interaction in the context of kidney transplantation. We use a mixed methods approach by combining a randomized design with a simple quantitative end point (frequency of conversations), different quantitative measurements for SDM, and several qualitative research methods (eg, records of physician-patient conversations and semistructured interviews) to examine the implementation of AI-based risk prediction in the clinic. Trial Registration ClinicalTrials.gov NCT06056518; https://clinicaltrials.gov/study/NCT06056518 International Registered Report Identifier (IRRID) PRR1-10.2196/54857
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.040 | 0.045 |
| Meta-epidemiology (narrow) | 0.008 | 0.003 |
| Meta-epidemiology (broad) | 0.014 | 0.008 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.003 | 0.005 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.009 | 0.009 |
| Insufficient payload (model declined to judge) | 0.054 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".