The Stroke Riskometer™ App: Validation of a Data Collection Tool and Stroke Risk Predictor
Bibliographic record
Abstract
BACKGROUND: The greatest potential to reduce the burden of stroke is by primary prevention of first-ever stroke, which constitutes three quarters of all stroke. In addition to population-wide prevention strategies (the 'mass' approach), the 'high risk' approach aims to identify individuals at risk of stroke and to modify their risk factors, and risk, accordingly. Current methods of assessing and modifying stroke risk are difficult to access and implement by the general population, amongst whom most future strokes will arise. To help reduce the burden of stroke on individuals and the population a new app, the Stroke Riskometer(TM) , has been developed. We aim to explore the validity of the app for predicting the risk of stroke compared with current best methods. METHODS: 752 stroke outcomes from a sample of 9501 individuals across three countries (New Zealand, Russia and the Netherlands) were utilized to investigate the performance of a novel stroke risk prediction tool algorithm (Stroke Riskometer(TM) ) compared with two established stroke risk score prediction algorithms (Framingham Stroke Risk Score [FSRS] and QStroke). We calculated the receiver operating characteristics (ROC) curves and area under the ROC curve (AUROC) with 95% confidence intervals, Harrels C-statistic and D-statistics for measure of discrimination, R(2) statistics to indicate level of variability accounted for by each prediction algorithm, the Hosmer-Lemeshow statistic for calibration, and the sensitivity and specificity of each algorithm. RESULTS: The Stroke Riskometer(TM) performed well against the FSRS five-year AUROC for both males (FSRS = 75.0% (95% CI 72.3%-77.6%), Stroke Riskometer(TM) = 74.0(95% CI 71.3%-76.7%) and females [FSRS = 70.3% (95% CI 67.9%-72.8%, Stroke Riskometer(TM) = 71.5% (95% CI 69.0%-73.9%)], and better than QStroke [males - 59.7% (95% CI 57.3%-62.0%) and comparable to females = 71.1% (95% CI 69.0%-73.1%)]. Discriminative ability of all algorithms was low (C-statistic ranging from 0.51-0.56, D-statistic ranging from 0.01-0.12). Hosmer-Lemeshow illustrated that all of the predicted risk scores were not well calibrated with the observed event data (P < 0.006). CONCLUSIONS: The Stroke Riskometer(TM) is comparable in performance for stroke prediction with FSRS and QStroke. All three algorithms performed equally poorly in predicting stroke events. The Stroke Riskometer(TM) will be continually developed and validated to address the need to improve the current stroke risk scoring systems to more accurately predict stroke, particularly by identifying robust ethnic/race ethnicity group and country specific risk factors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.041 | 0.081 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".