Predictors and associations of complications in ureteroscopy for stone disease using AI: outcomes from the FLEXOR registry
Bibliographic record
Abstract
We aimed to develop machine learning(ML) algorithms to evaluate complications of flexible ureteroscopy and laser lithotripsy(fURSL), providing a valid predictive model. 15 ML algorithms were trained on a large number fURSL data from > 6500 patients from the international FLEXOR database. fURSL complications included pelvicalyceal system(PCS) bleeding, ureteric/PCS injury, fever and sepsis. Pre-treatment characteristics served as input for ML training and testing. Correlation and logistic regression analysis were carried out by a multi-task neural network, while explainable AI was used for the predictive model. ML algorithms performed excellently. For intraoperative PCS bleeding, Extra Tree Classifier achieved the best accuracy at 95.03% (precision 80.99%), and greatest correlation with stone diameter(0.21) and residual fragments(0.26). PCS injury was best predicted by RandomForest (accuracy 97.72%, precision 63.50%). XGBoost performed best for ureteric injury (accuracy 96.88%, precision 60.67%). Both demonstrated moderate correlation with preoperative characteristics. Postoperative fever was predicted by Extra Tree Classifier with 91.34% accuracy (precision 58.20%). Cat Boost Classifier predicted postoperative sepsis with 99.15% accuracy (precision 66.38%), and the best overall performance. At logistic regression, postoperative fever/sepsis positively correlated with preoperative urine culture(p = 0.001). ML represents a powerful tool for automatic prediction of outcomes. Our study showed promises in algorithms training and validation on a very large database of patients treated for urolithiasis, with excellent accuracy for prediction of complications. With further research, reliable predictive nomograms could be created based on ML analysis, to serve as aid to urologists and patients in the decision making and treatment planning process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".