Commentary: Bringing down Babel—a pathway to a universal adverse events language
Bibliographic record
Abstract
Central MessageCan we build thoracic surgery programs that eliminate/minimize complications? Harmonizing adverse events classifications across databases is a key issue that first must be addressed.“Behold, the people is one, and they have all one language; and this they begin to do: and now nothing will be restrained from them, which they have imagined to do.”—Genesis 11:6See Article page 250 in the June 2021 issue. Can we build thoracic surgery programs that eliminate/minimize complications? Harmonizing adverse events classifications across databases is a key issue that first must be addressed. See Article page 250 in the June 2021 issue. Can we build a thoracic surgery program that eliminates or minimizes all complications? Is this a lofty height which we are destined never to reach? Complications after thoracic surgery have been decreasing, in part due to global efforts to objectively assess, document, and improve on outcomes.1Sigler G. Anstee C. Seely A.J.E. Harmonization of adverse events monitoring following thoracic surgery: pursuit of a common language and methodology.J Thorac Cardiovasc Surg Open. 2021; 6: 250-256Google Scholar While descriptions of postoperative complications in thoracic surgeries are widely reported, Sigler and colleagues1Sigler G. Anstee C. Seely A.J.E. Harmonization of adverse events monitoring following thoracic surgery: pursuit of a common language and methodology.J Thorac Cardiovasc Surg Open. 2021; 6: 250-256Google Scholar identify how the variety of systems used to classify adverse events (AE) undermine the potential of multicenter collaboration and data synthesis. In this issue of JTCVS Open, the authors describe their approach to harmonizing AE across databases.1Sigler G. Anstee C. Seely A.J.E. Harmonization of adverse events monitoring following thoracic surgery: pursuit of a common language and methodology.J Thorac Cardiovasc Surg Open. 2021; 6: 250-256Google Scholar The discordance between thoracic surgery AE databases has been previously described.2Ivanovic J. Seely A.J.E. Anstee C. Villeneuve P.J. Gilbert S. Maziak D.E. et al.Measuring surgical quality: comparison of postoperative adverse events with the American College of Surgeons NSQIP and the thoracic morbidity and mortality classification system.J Am Coll Surg. 2014; 218: 1024-1031Google Scholar,3Salati M. Refai M. Pompili C. Xiumè F. Sabbatini A. Brunelli A. Major morbidity after lung resection: a comparison between the European Society of Thoracic Surgeons Database System and the thoracic morbidity and mortality system.J Thorac Dis. 2013; 5: 217-222Google Scholar As such, it is difficult to draw valid comparisons between the AE of patients characterized using different databases. This precludes centers from pooling their outcomes data for meaningful international collaboration and quality improvement. However, the authors offer a means to collect AE in the same manner for their translation into any of the commonly used AE classification systems. A system is only as strong as the underlying assumptions on which it is constructed. Herein lies an important, but perhaps currently inescapable, methodologic limitation identified by the authors; the definitions of harmonization and the manner in which judgments are made regarding degree of harmonization are currently subjective and somewhat opaque. To harmonize differing definitions of AE, the definition of harmonization itself was characterized by a single author (if “perfect”) or by consensus with 2 authors (if not “perfect”). This subverts the reproducibility of harmonized definitions due to the necessarily subjective interpretation of AE elements. Certainly, there are definitions in each classification system that are straightforward or approximate objective measurement. However, inter-rater reliability cannot be taken for granted, especially in a novel undertaking predicated on the interpretation of a single expert. The authors acknowledge that it is good practice to ensure multidisciplinary discussion in efforts to characterize AE within individual institutions. Efforts to characterize AE across institutions should be held to a similar standard, if not more rigorous due to the risk of data being “lost in translation.” Therefore, such consensus-based discussions around definitions require not only standardized definitions of AE categories but also requisite thresholds of agreement among experts for an AE to be characterized. In doing so, one ensures a robust foundation on which to build this synthesis of multiple classifications. Nevertheless, a major strength of this paper is that it establishes a framework for harmonization. This creates the unique opportunity to compare AE data from institutions across the world. Currently, centers that desire to collaborate across multiple AE databases must fill in their AE data for each separate system. However, the approach of Sigler and colleagues1Sigler G. Anstee C. Seely A.J.E. Harmonization of adverse events monitoring following thoracic surgery: pursuit of a common language and methodology.J Thorac Cardiovasc Surg Open. 2021; 6: 250-256Google Scholar is simple, with 4 drop-down menus that facilitate AE classification on researchers' behalf. Therefore, researchers neither must enter the same data repeatedly nor learn an entirely new system to harmonize. Rather, the authors' approach aligns already-existing systems to harmonize. Because it is practical and user-friendly in this way, this system also ensures accessibility to researchers who seek to collaborate internationally for the first time. At the end of the day, authors should be commended for tackling an important problem in the way that we all communicate with each other, which is a barrier to sustainable and large-scale quality improvement initiatives across jurisdictions. Given the long-reaching implications that may result from such harmonization endeavors, it is imperative that underlying assumptions about definitions and decisions for harmonization be abundantly clear and reproducible. Otherwise, the very human frailty of subjectivity, rather than any vengeful deity, will be what brings down our proverbial tower before we reach the lofty heights of zero complications. Harmonization of adverse events monitoring following thoracic surgery: Pursuit of a common language and methodologyJTCVS OpenVol. 6PreviewThoracic surgery carries significant risk of postoperative adverse events (AEs). Multiple international recording systems are used to define and collect AEs following thoracic surgery procedures. We hypothesized that a simple-yet-ubiquitous approach to AE documentation could be developed to allow universal data entry into separate international databases. Full-Text PDF Open Access
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".