External validation of the DAYS score for suspected deep vein thrombosis
Bibliographic record
Abstract
Background: Diagnosing deep vein thrombosis (DVT) involves clinical assessment, D-dimer testing, and imaging. The DAYS score, a novel 2-item prediction tool combined with D-dimer, demonstrated promising performance but required external validation. Objectives: This study aimed to validate 2 DVT prediction scores: the DAYS score and a newly developed DVT score. Methods: Data were collected from a prospective Norwegian DVT management study (2015-2018; NCT02486445). The DAYS score includes 2 items (DVT most likely diagnosis and calf swelling), while the new score comprises tenderness along the deep veins and previous venous thromboembolism. DVT was considered excluded if no items were present and D-dimer was < 1.0 μg/mL or if ≥1 item was present and D-dimer was < 0.5 μg/mL. The 2-tier Wells score served as the reference. Safety was defined as the number of missed DVT cases divided by the total number of patients classified as having DVT excluded and was set at 2%. Results: Among 1312 patients (median age, 64 years; IQR, 52-73 years; 55% women), 261 (20.0%) had confirmed DVT. The DAYS score excluded DVT in 455 patients (34.6%), of whom 11 were diagnosed with DVT (failure rate, 2.4 %; 95% CI, 1.2-4.2). The new score excluded DVT in 519 patients (39.6%) and missed 7 cases with confirmed DVT (failure rate, 1.3%; 95% CI, 0.5-2.8). The Wells score excluded DVT in 271 patients (20.6%), missing only 2 cases with confirmed DVT. Conclusion: While both the DAYS score and the new score demonstrated low failure rates, they exceeded the predefined safety threshold.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.037 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".