Automated Pulmonary Embolism Risk Classification and Guideline Adherence for Computed Tomography Pulmonary Angiography Ordering
Bibliographic record
Abstract
BACKGROUND: The assessment of clinical guideline adherence for the evaluation of pulmonary embolism (PE) via computed tomography pulmonary angiography (CTPA) currently requires either labor-intensive, retrospective chart review or prospective collection of PE risk scores at the time of CTPA order. The recording of clinical data in a structured manner in the electronic health record (EHR) may make it possible to automate the calculation of a patient's PE risk classification and determine whether the CTPA order was guideline concordant. OBJECTIVES: The objective of this study was to measure the performance of automated, structured data-only versions of the Wells and revised Geneva risk scores in emergency department (ED) encounters during which a CTPA was ordered. The hypothesis was that such an automated method would classify a patient's PE risk with high accuracy compared to manual chart review. METHODS: We developed automated, structured data-only versions of the Wells and revised Geneva risk scores to classify 212 ED encounters during which a CTPA was performed as "PE likely" or "PE unlikely." We then combined these classifications with D-dimer ordering data to assess each encounter as guideline concordant or discordant. The accuracy of these automated classifications and assessments of guideline concordance were determined by comparing them to classifications and concordance based on the complete Wells and revised Geneva scores derived via abstractor manual chart review. RESULTS: The automatically derived Wells and revised Geneva risk classifications were 91.5 and 92% accurate compared to the manually determined classifications, respectively. There was no statistically significant difference between guideline adherence calculated by the automated scores compared to manual chart review (Wells, 70.8% vs. 75%, p = 0.33; revised Geneva, 65.6% vs. 66%, p = 0.92). CONCLUSION: The Wells and revised Geneva score risk classifications can be approximated with high accuracy using automated extraction of structured EHR data elements in patients who received a CTPA. Combining these automated scores with D-dimer ordering data allows for the automated assessment of clinical guideline adherence for CTPA ordering in the ED, without the burden of manual chart review.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.067 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".