Decision theoretical foundations of clinical practice guidelines: an extension of the ASH thrombophilia guidelines
Bibliographic record
Abstract
ABSTRACT: Decision analysis can play an essential role in informing practice guidelines. The American Society of Hematology (ASH) thrombophilia guidelines have made a significant step forward in demonstrating how decision modeling integrated within Grading of Recommendations Assessment, Developing, and Evaluation (GRADE) methodology can advance the field of guideline development. Although the ASH model was transparent and understandable, it does, however, suffer from certain limitations that may have generated potentially wrong recommendations. That is, the panel considered 2 models separately: after 3 to 6 months of index venous thromboembolism (VTE), the panel compared thrombophilia testing (A) vs discontinuing anticoagulants (B) and testing (A) vs recommending indefinite anticoagulation to all patients (C), instead of considering all relevant options simultaneously (A vs B vs C). Our study aimed to avoid what we refer to as the omitted choice bias by integrating 2 ASH models into a single unifying threshold decision model. We analyzed 6 ASH panel's recommendations related to the testing for thrombophilia in settings of "provoked" vs "unprovoked" VTE and low vs high bleeding risk (total 12 recommendations). Our model disagreed with the ASH guideline panels' recommendations in 4 of the 12 recommendations we considered. Considering all 3 options simultaneously, our model provided results that would have produced sounder recommendations for patient care. By revisiting the ASH guidelines methodology, we have not only improved the recommendations for thrombophilia but also provided a method that can be easily applied to other clinical problems and promises to improve the current guidelines' methodology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.165 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".