Clinical prediction tools for rare complications: are large administrative healthcare databases the answer?
Bibliographic record
Abstract
In this issue of Anaesthesia, Okada et al. harness information found in the Japanese Trauma Data Bank (JTDB), a large administrative database, to create a clinical prediction tool for emergency front-of-neck airway (eFONA) 1. The clinical predication tool, aptly named ‘eFONA’, is designed to assist emergency department physicians in anticipating the need for eFONA in a patient arriving from a trauma-related pre-hospital setting. A total of 198,182 trauma patients (between 2004 and 2017) from 231 Japanese hospitals were analysed. In total, 467 eFONA events were recorded representing 0.2% of all patients. This may very well be the largest data set of eFONA events recorded and their article is a valuable addition to our understanding of eFONA events in trauma. Although difficult to define, the concept of big data was first characterised by the three key concepts: high volume; variety; and velocity of data 2. Large, administrative healthcare data sets are considered a subset of ‘big data’ and offer unique opportunities to determine predictors of rare events such as eFONA. However, all big data are not created equally. Depending on the source, big data can be messy, incomplete and in a variety of formats. Big data derived from structured administrative healthcare databases represent a subset that typically contains defined variables that are collected for the purposes of heath programme administration and billing. Although the focus of this Editorial is on structured administrative healthcare databases, we refer interested readers to a previous editorial for a broader discussion of big data in airway research 3. Using administrative healthcare data for research purposes has several advantages. These types of data can be used to identify new patterns, trends and create predictive models for rare outcomes that would otherwise have been difficult to detect with traditional research methods and analytics 4. For many rare outcomes like eFONA, prospective randomised trials would be impractical, requiring an excessively large sample size 5. These databases offer large volumes of structured data with relatively defined variables, usually resulting in higher reliability and lower rates of missing data than other forms of big data. As these large databases are typically drawn from broad populations, they offer a valuable level of generalisability not observed in smaller clinical trials. Finally, administrative data can be cost-efficient and practical to retrieve when compared with prospective data collection. Unfortunately, healthcare database analysis also have a significant disadvantage; they are created and maintained to guide government policy and healthcare spending, not to answer specific research questions 6. This creates inherent limitations. Clinically-relevant variables may not be included in the database, which has important implications when attempting to apply these data to a clinical research question. We see these limitations in the article by Okada et al. Although the authors captured whether eFONA had occurred, this variable was collected as an umbrella term; specific details regarding the technique used and first-pass success rates were not collected and are unavailable. Similarly, patient details such as predictors of difficult mask-ventilation or tracheal intubation before their trauma are unavailable for analysis. This limitation is common to analyses of administrative data and restricts our ability to answer clinical questions. As will most analyses of administrative big data, the authors encountered missing data in their study. There are two main approaches to missing data in regression modelling, either through omitting these patients from the model (complete case analysis) or using an imputation strategy to estimate missing data points. Rather than solely conducting a complete case analysis which reduces the sample size, Okada et al. chose an imputation strategy by substituting missing values with the most frequent categorisation or the median value. Although this approach has the benefit of simplicity, it may result in artificial reduction in data variability by replacing all missing values with the most frequent data value 7, 8. This approach may also result in a biased estimate; however, this is less likely when the proportion of missing data is less than 10% 8-10. In the article by Okada et al., systolic blood pressure and level of consciousness were missing in 15.9% and 10.4%, respectively, of total patients. Some variables will be recorded in all patients, therefore, missing data can be quantified. Multiple imputation is an alternative technique which utilises multivariable regression to substitute missing variables. With the vast sample size and available predictors, this approach may be preferable. For other variables, such as presence of specific injuries (face/neck trauma) or mechanism of injury (motor bike/fall from a height), no total number is known. As a result, unlike basic baseline characteristics, potential missing data cannot be quantified. To ensure the accuracy and clinical utility of the prediction tool, validation in a new sample of patients is necessary. In study by Okada et al., the authors divided the total number of hospitals’ retrospective trauma patients into two groups: the prediction model cohort was developed from 116 hospitals (100,120 patients); and the validation cohort from 115 hospitals (98,062 patients). As hospitals, rather than patients, were randomly assigned to one of these two cohorts, this may introduce some bias into the study. For example, if a particular hospital is poor at recording data, missing data could be distributed differentially into one cohort in a non-random fashion. As seen in Table 1 provided by Okada et al., sex is missing in 105 patients in the validation cohort vs. 22 patients in the development cohort. Additionally, if patients are triaged to specific hospital(s) based on the requirement for specific speciality care (e.g. a neurosurgical centre), randomisation of hospitals will not balance these types of patients between the two cohorts. This can also be seen in Table 1, where the development cohort tended to be older, with more male patients than the validation cohort. These unresolved imbalances between the groups could impact the prediction tool's development and clinical utility; one potential solution is using patient-level random allocation. The unweighted clinical prediction tool presented by Okada et al. is a valuable contribution to our understanding of the need for eFONA in trauma; their tool now requires external validation. Eternal validation is especially important when the prediction model is for rare events. Model overfitting, where a model is fitted too closely to a specific set of data points, can severely limit the accuracy of a prediction model; this is more common when predicting rare events (e.g. eFONA) relative to the number of total events (total trauma patients) 11. In this context, the probability of an eFONA event tends to be underestimated in low-risk patients and overestimated in high-risk patients 11. Most prediction tools require further calibration in other populations and practice settings. For example, the performance of a clinical prediction tool can vary widely between specific patient populations, such as medical vs. surgical patients, or between surgical specialties 12. Previous analyses have suggested that the effect size estimates generated from large observational studies can be inflated and inaccurate 13, emphasising the need for prospective validation. Unfortunately, given the low incidence of eFONA in this study (33 per year in 231 hospitals or 0.15 eFONA events per hospital per year), prospective prediction rule studies in this area are likely to be impractical. Multiple country collaborative studies would be required to complete a validation study in a reasonable amount of time. At a minimum, however, assessing the performance of their tool in administrative databases in other countries should be an achievable next step. The relatively recent availability of big data in medical research has facilitated greatly the development of clinical prediction tools, particularly for rare events like eFONA. Unlike clinical trials that test the efficacy of an intervention and can directly inform clinical practice, the utility of clinical prediction tools can only determine associations and predictions. These tools have a distinct role for the clinician; they allow us to stratify our patient populations by their predicted risk and allocate resources accordingly. All acute care clinicians involved in airway management require the skill of performing an eFONA; however, most will only perform this once or twice in a career. Mental preparation, repeated deliberate practice, equipment and team simulation are required to maintain eFONA skills. Delay in both the recognition of the need for an eFONA procedure, and delay procedure itself, has been a recurrent theme in closed claims analysis 14. The eFONA prediction tool may allow the clinician and team to mentally prepare for a higher risk patient. As a result of big data, large randomised trials and meta-analyses, clinical prediction tools are intuitively appealing as they may aid in efficiently managing a myriad of clinical conditions. However, caution is required before adoption of a clinical decision rule unless the tool has been validated in the approximate population a clinician would manage. Even better, a clinical prediction tool that has then been shown to benefit patient-centred outcomes, such as prevention of hypoxic brain damage or death, is the best level of evidence for a clinical predication tool for rare events such as eFONA 15. Overall, the authors should be congratulated for their work in capturing administrative data on a large scale to assist clinicians in determining the risk of eFONA before patient arrival in the Emergency Department. As expected, the model is imperfect and subject to many of the inherent limitations of using administrative healthcare big data sets. However, the rarity of eFONA makes clinical trials and prospective trials impractical. The eFONA clinical prediction tool provides a novel and useful set of criteria to preparation for such an event. As the use of healthcare databases to create clinical predication tools will continue to grow, a multi-step process is required to validate and adopt the right clinical predication tools for the right patient populations for the most important patient-centred outcomes. The creation of consensus guidelines for the use, reporting and analytic approaches to healthcare database research in the creation of clinical predication tools 15, similar to established guidelines for systematic reviews 16 or randomised controlled trials 17, would be a valuable step forward in harnessing the power of big data. LD is an Editor of Anaesthesia. No other competing interests declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".