Developing a Diagnostic Multivariable Prediction Model for Urinary Tract Cancer in Patients Referred with Haematuria: Results from the IDENTIFY Collaborative Study
Bibliographic record
Abstract
BACKGROUND: Patient factors associated with urinary tract cancer can be used to risk stratify patients referred with haematuria, prioritising those with a higher risk of cancer for prompt investigation. OBJECTIVE: To develop a prediction model for urinary tract cancer in patients referred with haematuria. DESIGN, SETTING, AND PARTICIPANTS: A prospective observational study was conducted in 10 282 patients from 110 hospitals across 26 countries, aged ≥16 yr and referred to secondary care with haematuria. Patients with a known or previous urological malignancy were excluded. OUTCOME MEASUREMENTS AND STATISTICAL ANALYSIS: The primary outcomes were the presence or absence of urinary tract cancer (bladder cancer, upper tract urothelial cancer [UTUC], and renal cancer). Mixed-effect multivariable logistic regression was performed with site and country as random effects and clinically important patient-level candidate predictors, chosen a priori, as fixed effects. Predictors were selected primarily using clinical reasoning, in addition to backward stepwise selection. Calibration and discrimination were calculated, and bootstrap validation was performed to calculate optimism. RESULTS AND LIMITATIONS: The unadjusted prevalence was 17.2% (n = 1763) for bladder cancer, 1.20% (n = 123) for UTUC, and 1.00% (n = 103) for renal cancer. The final model included predictors of increased risk (visible haematuria, age, smoking history, male sex, and family history) and reduced risk (previous haematuria investigations, urinary tract infection, dysuria/suprapubic pain, anticoagulation, catheter use, and previous pelvic radiotherapy). The area under the receiver operating characteristic curve of the final model was 0.86 (95% confidence interval 0.85-0.87). The model is limited to patients without previous urological malignancy. CONCLUSIONS: This cancer prediction model is the first to consider established and novel urinary tract cancer diagnostic markers. It can be used in secondary care for risk stratifying patients and aid the clinician's decision-making process in prioritising patients for investigation. PATIENT SUMMARY: We have developed a tool that uses a person's characteristics to determine the risk of cancer if that person develops blood in the urine (haematuria). This can be used to help prioritise patients for further investigation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.036 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".