Bibliographic record
Abstract
Dr Dekker and colleagues assert that prediction is difficult, and recommend that researchers focus efforts on validation and implementation studies rather than the development of new models. They argue that flaws in model development and questionable findings from impact studies have limited the clinical utility of most risk prediction models. We would agree with Dr Dekker on the importance of external validation and studies of clinical impact. For predicting kidney failure, the field certainly needs to move on to evaluate the utility of the KFRE, rather than its discriminatory performance. However, we think significant deficits still remain for model development and early validation for predicting cardiovascular disease and early mortality on dialysis. In our systematic review in 2013, we demonstrated a lack of accurate models for predicting cardiovascular events in patients with CKD, and recent reviews suggest a similar lack for predicting early mortality in dialysis [1]. In the absence of these models, clinicians may choose high-dose statins or other more expensive cardioprotective agents without knowing their patient’s absolute risk, or may recommend dialysis without knowing the expected survival benefit [2]. This lack of objective information can lead to impaired shared decision-making, and often results in paternalistic decisions, leaving patients less empowered and more dependent [3]. Clinicians and researchers should continue to develop new models for cardiovascular disease in CKD, and early mortality in dialysis. These models should be externally validated, easy to use and tied to actionable thresholds for clinical decision-making. A framework for model development that emphasizes the importance of the entire process, from cohort and variable selection to knowledge translation using patients and physicians at the bedside, should be followed. Researchers, policy makers, clinicians and patients can work together in this framework to generate truly useful risk prediction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.190 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.007 | 0.005 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.023 | 0.023 |
| Insufficient payload (model declined to judge) | 0.129 | 0.049 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".