Bibliographic record
Abstract
Adverse outcomes have been a matter of concern from the earliest attempts at transfusion and there have been continuous efforts to ameliorate them—perhaps best exemplified by the management of posttransfusion hepatitis over the past 60 years.1 However, it was not until the emergence of HIV/AIDS that there were organized attempts to measure the frequency of adverse outcomes on a national scale—an effort originating in France and the United Kingdom and now known as hemovigilance.2, 3 The ultimate objective of these programs is to use surveillance to inform priority setting and to facilitate measures to control the adverse events. These approaches have demonstrated success in reducing the frequency of a number of serious outcomes, such as transfusion-related acute lung injury (TRALI). Indeed, in this case, interventions developed and validated as a result of hemovigilance programs have subsequently been implemented even in countries without such monitoring, for example, the use of predominantly male plasma in the United States. As with all surveillance programs, hemovigilance should have rigorous requirements for consistency. Case definitions must be precise and must be consistently used by all those who report results. For example, it is not easy to differentiate TRALI from transfusion-associated circulatory overload (TACO), but their etiology and hence the necessary interventions differ greatly. Care must be taken to assure that imputability is accurately determined and it is critical that denominators are uniformly reported. These issues place a premium on the quality of the surveillance process and demand validation and continuing evaluation of programs. In this context, the French, Canadian, and British systems, along with a number of others,4 have mechanisms in place to review and, if necessary, reclassify reports of serious adverse outcomes before they are accepted as part of the record. This set of procedures usually involves discussion of case details with the reporting institution. Such a process is, of course, resource-intensive and may best be achieved in a unified, government-managed medical care system. The US hemovigilance system is in its infancy and, by comparison with other programs, is less robustly controlled,5 in part due to the nature of the US health care delivery system. It is voluntary and no resources external to the participating hospitals have been provided for validation of input data. The system relies on unmonitored adherence to detailed case definitions after training of those responsible for reporting, along with written procedure manuals. Reports are submitted to the US Centers for Disease Control and Prevention (CDC) through a module of the National Healthcare Safety Network (NHSN). The data are managed under conditions that assure confidentiality for reporting institutions, and there is no mechanism at the CDC to directly review the submitted data with those providing the reports. However, the system was developed jointly by AABB and CDC and the AABB simultaneously established a patient safety organization (Center for Patient Safety [CPS]) in which hospitals reporting to NHSN may participate. Such participation explicitly allows the CPS to receive and review attributable data submitted by that hospital to the NHSN. This mechanism has become an invaluable tool to allow the AABB to partner with the CPS member hospitals. More specifically, it does offer an avenue for retrospective review of the validity of reported data. In this issue of TRANSFUSION, Harvey and her colleagues6 from the CDC review the first 3 years of data reported to the NHSN Hemovigilance Module, covering just over 2.1 million blood components transfused, with 5136 reported adverse reactions, for a rate of 239.5 reactions per 100,000 components. The data came from 77 transfusing facilities and were estimated to reflect 2.1% to 4.5% of components transfused nationally, depending on the year. Encouragingly, the reported rates were not greatly different from those seen in other major hemovigilance systems. However, the authors note some deficiencies in the system in addition to the absence of verification of data, which will be discussed below. The data have not been subjected to any statistical analysis, in part because the authors do not consider them to be nationally representative and thus they are not considered to be a legitimate sample. There is thus no way to account realistically for variability between locations. Of additional concern is the fact that the treatment of denominators is unclear in the context of pooled products, where a therapeutic dose is composed of multiple components, or for apheresis products, where a single donation may be used to treat two or more patients. Thus, the apparently increased reaction rates from apheresis compared to whole blood–derived platelets (PLTs) require clarification, as does the smaller difference between red blood cells from apheresis versus those from whole blood. The difference between leukoreduced and nonleukoreduced PLTs is most likely due to the fact that apheresis PLTs are leukoreduced in preparation, while many whole blood–derived PLTs are issued without leukoreduction, but this issue will probably require multivariate statistical analysis for resolution. These concerns may eventually be managed within the existing system, but it is more probable that the absence of a data validation step will continue to confound the interpretation of data from the system. If so, what should be done? Some data suggest that this question is more than academic. Recently, AuBuchon and colleagues7 reported on a study in which a number of institutions were provided with hypothetical reports of 36 cases of transfusion reactions and were asked how they would report them to the NHSN Hemovigilance Module. The responses were compared to those of an expert group. Overall, there was a 72.1% agreement on the adverse event type, a 76.5% match with the case definition, a 69.6% agreement on severity, and a 64.4% agreement on imputability. Interestingly, overall rates did not differ between institutions that had or had not received the standard training for the NHSN program. The respondents had most difficulty with classifying transfusion-associated dyspnea (TAD) and TACO. Paradoxically, untrained participants appeared to be more accurate in reporting transfusion-transmitted infection than did those who had been trained. A working group established by AABB took a different approach by asking selected participants in the AABB's CPS to provide supplementary clinical and diagnostic information relating to cases of pulmonary and infectious reactions that had been reported through the NHSN. A small expert group then reviewed the data to determine the accuracy of reporting. The authors of this Editorial are members of the working group (see listing at end of article). The group examined 14 reports of transfusion-transmitted infections. All were bacterial infections. Five of these, one of which was reported as definitive, did not meet the case definition inasmuch as there was no evidence of an organism in the patient. The other nine cases were both reported and adjudicated by the experts as definitive. With respect to imputability, four cases (including the one that did not meet the case definition) were reported as definitive, but the expert panel agreed in only one case and there was uncertainty about the imputability of two of them. Of the remainder, one case was reported as having probable imputability, but the panel judged it to be definitive. The remainders were reported as “probable” or “possible,” while the panel considered them to be “possible” or “doubtful.” The situation is more disconcerting for the 42 pulmonary cases reported. Of six “definitive” TRALI reports, the panel agreed on only one. Two were considered to be possible, one was judged to be TACO, and one to be TAD, while the sixth did not meet the case definitions for respiratory reactions. In only one case did the panel agree on the reported imputability (possible): the case reported as definitive should have been reported as TAD. A number of cases were reported as “probable,” a grade that was not available for TRALI cases at that time. For TACO, 27 cases were reported, but the panel agreed with only four case definitions and on examination of supplementary data, the panel agreed that there were insufficient data to support a report of TACO for 20 of the cases. One case reported as definitive TACO should have been reported as TRALI. Again, imputability was unclear in many of the cases. Pulmonary complications are notoriously difficult to diagnose at the bedside and are often accompanied by confounding comorbidities. The panel members themselves often found that they were actually trying to make a diagnosis, rather than to determine whether the data supported the strict case definition. We suspect that this tendency, while understandable, contributes to the variability and some of the inaccuracies noted in the program. This issue was also identified by AuBuchon and colleagues. It is a subtle issue, but it impacts the consistency and therefore the value of a surveillance system, which is critically dependent on adhering to case definitions. As we recognize above, the US hemovigilance system is in its infancy. It is reasonable to believe that, as was the case with the hospital infections component of NHSN, improvement will occur over time if enough resources are committed and enough effort is expended.8 However, it should also be noted that training and education are probably not adequate to assure that the current system will be improved.7 How likely is it that resources will become available to implement a data verification loop in the existing system? Additionally, the current system does not have any internal verification mechanisms: for example, our reviews revealed that a case of TRALI can be (and has been) reported without a patient X-ray having been performed: this omission could not be determined from the NHSN report and required access to supplementary information through the CPS. As suggested by AuBuchon, the use of automated algorithms to report and classify reactions could help to reduce human bias in the system.7 Another question to ponder is whether our objective is too broad. Every year, more that 15 million components are transfused in the United States. Could we not get good information by careful scrutiny of a representative sample—less quantity, more quality? As a result of our concerns, the AABB is already collecting supplementary data on serious adverse events from some of the members of the CPS. These data can be used retrospectively to validate and/or correct data from a subset of the NHSN reports, provided that the necessary resources can be maintained. Finally, it must be noted that the American Red Cross and Blood Systems, Inc., for example, have developed and maintained their own vigilance systems and have used their own data to effect and demonstrate the efficacy of changes in bacterial safety measures and management of TRALI. We recommend that, at a minimum, funding to support appropriate automation and routine validation of reports of serious adverse events to NHSN and/or the CPS be established and maintained. Members of the AABB Working Group on Quality and Validation are: B. Custer, R. Dodd, A. Eder, M. Fung, L. Katz, S. Kleinman, J. Malasky, E. Notari, P. Robillard, and B. Whitaker. The authors have disclosed no conflicts of interest. Roger Y. Dodd, PhD1 ([email protected]) Louis M. Katz, MD2 1North Potomac, MD 2America's Blood Centers Washington, DC
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.070 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".