Predictive Modeling of Urinary Stone Composition Using Machine Learning and Clinical Data: Implications for Treatment Strategies and Pathophysiological Insights
Bibliographic record
Abstract
Purpose: Preventative strategies and surgical treatments for urolithiasis depend on stone composition. However, stone composition is often unknown until the stone is passed or surgically managed. Given that stone composition likely reflects the physiological parameters during its formation, we used clinical data from stone formers to predict stone composition. Materials and Methods: Data on stone composition, 24-hour urine, serum biochemistry, patient demographics, and medical history were prospectively collected from 777 kidney stone patients. Data were used to train gradient boosted machine and logistic regression models to distinguish calcium vs noncalcium, calcium oxalate monohydrate vs dihydrate, and calcium oxalate vs calcium phosphate vs uric acid stone types. Model performance was evaluated using the kappa score, and the influence of each predictor variable was assessed. Results: The calcium vs noncalcium model differentiated stone types with a kappa of 0.5231. The most influential predictors were 24-hour urine calcium, blood urate, and phosphate. The calcium oxalate monohydrate vs dihydrate model is the first of its kind and could discriminate stone types with a kappa of 0.2042. The key predictors were 24-hour urine urea, calcium, and oxalate. The multiclass model had a kappa of 0.3023 and the top predictors were age and 24-hour urine calcium and creatinine. Conclusions: Clinical data can be leveraged with machine learning algorithms to predict stone composition, which may help urologists determine stone type and guide their management plan before stone treatment. Investigating the most influential predictors of each classifier may improve the understanding of key clinical features of urolithiasis and shed light on pathophysiology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".