Deployment of a human-centred clinical decision support system for pulmonary embolism: evaluation of impact on quality of diagnostic decisions
Bibliographic record
Abstract
Pulmonary embolism (PE) is a serious condition that presents a diagnostic challenge for which diagnostic errors often happen. The literature suggests that a gap remains between PE diagnostic guidelines and adherence in healthcare practice. While system-level decision support tools exist, the clinical impact of a human-centred design (HCD) approach of PE diagnostic tool design is unknown. DESIGN: Before-after (with a preintervention period as non-concurrent control) design study. SETTING: Inpatient units at two tertiary care hospitals. PARTICIPANTS: General internal medicine physicians and their patients who underwent PE workups. INTERVENTION: After a 6-month preintervention period, a clinical decision support system (CDSS) for diagnosis of PE was deployed and evaluated over 6 months. A CDSS technical testing phase separated the two time periods. MEASUREMENTS: PE workups were identified in both the preintervention and CDSS intervention phases, and data were collected from medical charts. Physician reviewers assessed workup summaries (blinded to the study period) to determine adherence to evidence-based recommendations. Adherence to recommendations was quantified with a score ranging from 0 to 1.0 (the primary study outcome). Diagnostic tests ordered for PE workups were the secondary outcomes of interest. RESULTS: Overall adherence to diagnostic pathways was 0.63 in the CDSS intervention phase versus 0.60 in the preintervention phase (p=0.18), with fewer workups in the CDSS intervention phase having very low adherence scores. Further, adherence was significantly higher when PE workups included the Wells prediction rule (median adherence score=0.76 vs 0.59, p=0.002). This difference was even more pronounced when the analysis was limited to the CDSS intervention phase only (median adherence score=0.80 when Wells was used vs 0.60 when Wells was not used, p=0.001). For secondary outcomes, using both the D-dimer blood test (42.9% vs 55.7%, p=0.014) and CT pulmonary angiogram imaging (61.9% vs 75.4%, p=0.005) was lower during the CDSS intervention phase. CONCLUSION: A clinical decision support intervention with an HCD improves some aspects of the diagnostic decision, such as the selection of diagnostic tests and the use of the Wells probabilistic prediction rule for PE.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".