Appropriateness of diagnostic strategies for evaluating suspected venous thromboembolism
Bibliographic record
Abstract
It was the objective of this study to determine the proportion of patients who undergo an appropriate diagnostic work-up following a D-dimer test performed to evaluate suspected pulmonary embolism (PE) or deep vein thrombosis (DVT). We performed a retrospective cohort study at a tertiary care hospital. We included patients if they underwent D-dimer testing between 2002 and 2005, if the D-dimer was performed for evaluation of VTE, and if the D-dimer test was successful. We classified: the patients' clinical probability of DVT or PE according to the Wells models, the imaging results, and the appropriateness of the testing algorithm. Of 1,000 randomly selected patients, 863 met our study criteria. Seven hundred nineteen patients (83%) had testing during an emergency department visit, while 144 were tested as inpatients (17%). Physicians performed the D-dimer test to evaluate DVT and PE in 238 (28%) and 625 (72%) patients, respectively. Overall, the testing strategy was appropriate in 69% (95% confidence interval [CI]: 66%-72%) of cases. The testing strategy was more likely to be appropriate for emergency department versus inpatients (75% vs. 39%, p < 0.05) and for DVT versus PE patients (84% vs. 63%, p < 0.05). Of all in-appropriately tested patients, under-utilization of diagnostic imaging was more common than over-utilization (90% vs. 10%, p < 0.05). VTE was confirmed in 37 of 138 'DVT patients' and 35 of 625 'PE patients' (16% [95% CI: 11%-21%] and 6% [95% CI: 4%-8%], respectively). In conclusion, physicians often fail to use diagnostic testing strategies for VTE correctly following a D-dimer test.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".