Bibliographic record
Abstract
In this issue of Haemophilia, Sharma and colleagues describe a quality improvement initiative to introduce a bleeding assessment tool (BAT), specifically a self-administered (rather than expert-administered) electronic BAT (e self-BAT), into practice at their well-known hematology program. Using Plan-Do-Study-Act methodology, they were able to get physicians to document BAT scores routinely, and to have the majority of referred patients complete the e self-BAT before their clinic visit. Using this abundance of BAT scores, they showed that the e self-BAT has moderate concordance with an expert-administered BAT, and then evaluated the utility of the e self-BAT in detecting patients who do—or do not—have a congenital bleeding disorder. The role of BATs in clinical practice is best established in the investigation of von Willebrand disease (VWD), as is discussed in a recent American Society of Hematology guideline.1 BATs are recommended in practice settings where the probability of VWD is low, as a standardized method for eliminating a need for laboratory testing in patients who do not have significant bleeding symptoms. The ASH guideline recommends against using BATs to exclude a need for laboratory testing in patients who have a higher probability of VWD, such as those who have been referred to a hematologist or those with an affected first-degree relative. (The ASH guideline also observes that, even in settings where their use in screening out patients who do not require laboratory testing is not appropriate, BATs can function as a standardized method for documenting the severity of bleeding symptoms, and can be used in this manner as part of an initial diagnostic process). The form of these recommendations from ASH are important: the utility of BATs is considered for a particular purpose (assisting decisions about performing diagnostic testing) regarding a specific disorder (VWD), and with reference to practice setting (as a proxy for information about epidemiology). For other purposes, BATs may not perform well. For example, in a study that evaluated the ability of the pediatric bleeding questionnaire (PBQ) to predict operative bleeding in patients undergoing elective surgeries, the PBQ had a sensitivity and positive predictive value (PPV) of 0%, albeit with a negative predictive value (NPV) of 98%.2 So, while Sharma and colleagues demonstrate that a physician or a patient can be made to complete a BAT, a question remains: what is going to be done with that BAT score? Following the success of BATs in informing diagnostic workups for VWD, Sharma et al. put the e self-BAT to work in diagnosing bleeding disorders more generally. There is some support for extending the use of BATs outside the specific context of VWD, although this support is mixed: for example, one study found that a BAT had better specificity for inherited platelet function disorders than for type 1 VWD, but worse sensitivity for inherited thrombocytopenias than for type 1 VWD.3 In the present study, the e self-BAT, with a PPV of 25% and an NPV of 74%, performs reasonably well. Or does it? In their study cohort of 79 patients, 22 (or 28%) had a laboratory-defined bleeding disorder. This means that a completely unintelligent diagnostic process, for example randomly guessing or deciding a priori that no patient had a bleeding disorder, would have an NPV of 72% ( = 100% − 28%). Most of the NPV is provided by the rarity of the disorders under consideration, and a tool such as a BAT needs to work hard to contribute additional value. This is not a unique finding in studies of BATs. In the PBQ study discussed above, for example, only 1 of 60 patients, approximately 2%, had an operative bleeding complication2: the 98% NPV in this study occurred essentially for free. In a prior study of the PBQ, 6 of 151 children (4%) met laboratory criteria for VWD, and the PBQ's negative predictive value was 99%,4 only slightly greater than the minimum of 96% implied by the study population. BATs may do better at detection than exclusion. The PBQ, for example, has a sensitivity of 83%; in the general pediatrics population in which the tool was first evaluated, with a lower prevalence of VWD than what would be expected in a hematologist's clinic, the PPV was 14%,4 which is likely high enough to warrant laboratory testing in most hematologists’ minds. Sharma provides a suggested algorithm for the incorporation of BATs into practice: those patients with a negative or normal BAT score would receive only an entry-level workup aimed at detecting thrombocytopenia, VWD, and those coagulation disorders severe enough to prolong routine coagulation times. More extensive workup, for platelet function disorders or rare factor deficiencies, would be offered to patients with elevated BAT scores, family histories of bleeding, or according to the clinician's discretion. This practice supposes a few things. The first is that the PPV of the e self-BAT in this practice setting (25%) is sufficient to warrant testing. This seems fair to say. But more thought is required regarding the proposed response to a negative BAT score: the NPV of the e self-BAT (74%) is not high enough to eliminate a need for testing. But what testing should occur? The use of a BAT to assign patients to different tiers of investigational effort is not yet supported by robust evidence, particularly with regard to rarer bleeding disorders. Determining the role—or roles—of BATs in hematology practice, then, requires continued investigation. In what practice settings is it appropriate to use BATs to determine a need for laboratory testing? And for which disorders? When do BAT scores predict future bleeding behaviour? Can BAT scores predict efficacy of treatment? There is much research to be done. And if BATs, as currently constructed, do not function reliably in these respects, it may be that new tools are required. The findings of Sharma et al. are useful in justifying what seems, anecdotally at least, to be a common practice among hematologists: ordering blood work on all patients referred for suspicion of a bleeding disorder, regardless of how severe those symptoms turn out to be when assessed by a bleeding disorders expert. These findings should also remind us of something that is often forgotten in clinical epidemiology: while sensitivity and specificity are properties of a test, PPV and NPV are properties of a test in a particular population. If the population changes, if the epidemiology of the phenomenon of interest changes, then the performance of the test changes, perhaps in ways that should make us carefully examine the proposed utility of the test. The author declares no conflicts of interest. Data sharing not applicable to this article as no datasets were generated or analysed during the current study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.025 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.008 | 0.007 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.010 | 0.016 |
| Insufficient payload (model declined to judge) | 0.016 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".