Derivation of a Simple Clinical Model to Categorize Patients Probability of Pulmonary Embolism: Increasing the Models Utility with the SimpliRED D-dimer
Bibliographic record
Abstract
We have previously demonstrated that a clinical model can be safely used in a management strategy in patients with suspected pulmonary embolism (PE). We sought to simplify the clinical model and determine a scoring system, that when combined with D-dimer results, would safely exclude PE without the need for other tests, in a large proportion of patients. We used a randomly selected sample of 80% of the patients that participated in a prospective cohort study of patients with suspected PE to perform a logistic regression analysis on 40 clinical variables to create a simple clinical prediction rule. Cut points on the new rule were determined to create two scoring systems. In the first scoring system patients were classified as having low, moderate and high probability of PE with the proportions being similar to those determined in our original study. The second system was designed to create two categories, PE likely and unlikely. The goal in the latter was that PE unlikely patients with a negative D-dimer result would have PE in less than 2% of cases. The proportion of patients with PE in each category was determined overall and according to a positive or negative SimpliRED D-dimer result. After these determinations we applied the models to the remaining 20% of patients as a validation of the results. The following seven variables and assigned scores (in brackets) were included in the clinical prediction rule: Clinical symptoms of DVT (3.0), no alternative diagnosis (3.0), heart rate >100 (1.5), immobilization or surgery in the previous four weeks (1.5), previous DVT/PE (1.5), hemoptysis (1.0) and malignancy (1.0). Patients were considered low probability if the score was <2.0, moderate of the score was 2.0 to 6.0 and high if the score was over 6.0. Pulmonary embolism unlikely was assigned to patients with scores < or =4.0 and PE likely if the score was >4.0. 7.8% of patients with scores of less than or equal to 4 had PE but if the D-dimer was negative in these patients the rate of PE was only 2.2% (95% CI = 1.0% to 4.0%) in the derivation set and 1.7% in the validation set. Importantly this combination occurred in 46% of our study patients. A score of <2.0 and a negative D-dimer results in a PE rate of 1.5% (95% CI = 0.4% to 3.7%) in the derivation set and 2.7% (95% CI = 0.3% to 9.0%) in the validation set and only occurred in 29% of patients. The combination of a score < or =4.0 by our simple clinical prediction rule and a negative SimpliRED D-Dimer result may safely exclude PE in a large proportion of patients with suspected PE.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.024 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".