Estimating the Risk of a Rare But Plausible Complication That Has Not Occurred After n Trials
Bibliographic record
Abstract
For any treatment to become acceptable, its risk-benefit ratio should be comparable or superior to alternative treatments. Quantifying risks is therefore important to regulatory authorities and in clinical decision making. The special case of risk estimation when a complication of interest has had 0 occurrences after n trials presents a special challenge. A hospital installs a modern air filtration system in the arthroplasty room and enforces an infection prevention bundle. After 50 arthroplasties, there was not 1 case of wound infection. Can it declare the risk of prosthesis infection at the facility 0? In a pilot project using a computer-controlled system for sedation of endoscopy patients, no intervention by the standby anesthesiologist was needed in the first 100 cases. Is the automatic system safe enough without immediate anesthesia backup? A common and intuitive way of determining the risk of a complication associated with an intervention is by dividing the number of times the complication has occurred by the number of interventions/trials performed. Statisticians call this the maximum likelihood estimate. But how does one estimate the risk when the plausible complication is rare and has not (yet) occurred? When the numerator is 0, risk estimation can still be done. For example, in the earlier arthroplasty example, the hospital could use a Bayesian approach by incorporating the arthroplasty infection rate before its new infection control strategy or the average infection rate of other facilities that have implemented a similar strategy in determining their poststrategy risk after n uneventful local cases. In 1983, Hanley and Lippman-Hand1 derived the “Rule of 3,” which states that if no complication has occurred after an intervention has been trialed ≥30 times, one can approximate the 1-tailed 95% CI of the risk to be (0–3/n), where n is the number of trials. This rule is derived as follows1,2: Let the risk of the plausible complication of interest occurring be r. The chance of the complication not occurring is (1 − r). The chance of the complication not occurring after n trials is (1 − r)n. If we are 100% confident that no complication should occur after another n trials, then (1 − r)n = 1, which means r = 0. This would imply that joint infections never occur after arthroplasty and computers for automatic sedation never crash. If the complication is serious, then we should not be so sure that we would again get 0 complications with the next n trials. How about 50% sure (ie, (1 − r)n = 0.5)? It is probably not good enough to tell patients the risk estimate is a coin toss and that the true risk could be higher or lower. In clinical trials, a P value >.05 would not allow investigators to declare statistical significance. We should exercise the same caution here. We already know that equating (1 − r)n to 1 gives a risk estimate of 0, and equating (1 − r)n to values close to 1 gives a risk estimate of nearly 0 (both potentially reckless). Equating (1 − r)n to 0.05, in contrast, produces a conservative r value that reflects a nearly worst-case scenario, and the chance that the true value of r would be even higher is ≤5%. In statistical parlance, if the true risk is higher than the r value derived from (1 − r)n = 0.05, then the likelihood of obtaining 0 complications after n trials is ≤5%. Solving (1 − r)n = 0.05, . If n is large, then the second and subsequent fractions of this arithmetic series rapidly diminish (denominators increase exponentially, numerators increase by a factor of 3) and r ≈ 3/n. For example, for n = 30, the second and third fractions are 5% and 0.17% that of the first fraction, respectively. If n = 60, then these numbers become 2.5% and 0.04%, respectively. Because there is a minus sign preceding the second fraction (which the Rule of 3 ignores), 3/n will always overestimate the upper limit of the 1-sided 95% CI based on the arithmetic series. As shown previously, for n = 30, it is ~5%, and for n = 60, it is ~2.5%. For a smaller n, the error is larger, thus the recommendation that the Rule of 3 should not be applied when n < 30. In statistics, when dealing with discrete outcomes (complication or no complication, success or failure), the probability of a certain outcome follows a binomial distribution. There are many ways of calculating CIs. Some readers will recognize that the solution for the equation (1 − r)n = 0.05 is the upper limit of a 2-tailed 90% CI of a binomial distribution. The result is identical to that derived from the commonly used Clopper-Pearson Exact method, which gives a conservative estimate and is, depending on the complication, sometimes preferred by regulatory agencies. There are approximate methods, generally not quite as conservative as the Clopper-Pearson Exact method and less accurate at close to the probability extremes (rare or common events), of calculating CIs, some of which are presented in Table 1. The differences between the different methods are probably of more interest to statisticians and may or may not be large enough to change the clinical perspective or the scientific analysis of the problem at hand.Table 1.: Estimating the Risk of Spinal Hematoma in Thrombocytopenic Parturients Receiving Neuraxial Block, Based on n Trials With 0 Occurrences of Epidural HematomaIn anesthesiology, the fact that many dreaded complications are (thankfully) rare often makes risk estimation challenging. An example is whether it is safe to use inhalational induction in infantile pyloromyotomy. Given the low incidence of aspiration, the potential for cricoid pressure (CP) to hamper airway management, the lack of randomized controlled trials, and so on, some anesthesiologists have challenged the continuation of the practice.4 Whereas CP is still broadly applied in the United States,5 several European guidelines no longer endorse CP.6–8 In 2001, Engelhardt et al9 reported that ≤50% of pediatric anesthesiologists did not use CP in situations where precautions for a “full stomach” would normally be taken. A 2009 survey showed that one-third of German anesthesiologists never use CP in conditions commonly accepted as an indication for CP.10 A British group11 has injected much needed quantitative data into the controversy by reporting 252 uncomplicated inhalational inductions for infantile pyloromyotomy. Sofer et al12 have countered that, based on that series (and without knowing the number of unreported cases of aspiration and the extent of how often inhalational induction is used worldwide), the risk of aspiration from inhalational induction in infants with pyloric stenosis is not 0 but ≤3/252 = 1.2%, stated with 95% confidence. Ho et al2 used the same principle to estimate the risk of neuraxial hematoma associated with thoracic epidural catheterization in cardiac surgery. Similarly, Lee et al3 estimated the risk of neuraxial hematoma in parturients with thrombocytopenia receiving an epidural (Table 1). They reviewed the Multicenter Perioperative Outcomes Group (MPOG) database, which consists of data from >50 hospitals across 17 states and in the Netherlands, to collate the number of epidural cases reported in 3 ranges of thrombocytopenia (<49,000, 29,000–69,000, and 70,000–99,000 mm−3). They also performed a systematic review and combined the MPOG data with summaries from the literature (Table 1). They found no case of clinically significant spinal hematoma.3 To illustrate our point earlier regarding how the Rule of 3 overestimates the risk calculated using the arithmetic series, notice that Lee et al3 overestimate the actual 1-tailed 95% CI upper limit by 11% for n = 15 and by 4.2% for n = 36 (Table 1). Websites (eg, https://www.danielsoper.com/statcalc/calculator.aspx?id=85, http://statpages.info/confint.html, http://epitools.ausvet.com.au/content.php?page=CIProportion) provide a place where users can simply input the number of complications/failures and the number of trials to obtain the desired CI (Table 1). So far, we have rightfully focused on the upper limit of the estimated risk (near worst-case scenario). However, the lower limit, taken to be 0 when there have been 0 occurrences, should not be completely ignored. Imagine citing the 95% CI estimate of epidural hematoma to a patient with preeclampsia who is in pain and has a platelet count of 55,000 mm−3 as 0%–3% (Table 1). The anesthesiologist can state that the 3% chance of neuraxial hematoma is highly conservative and unlikely to be exceeded, and the good news is that the risk can be as low as 0, and an epidural may help to normalize her blood pressure. We respectfully submit that while a quote of 0 risk may be understood by a statistician, it is misleading to a layperson, especially someone in distress desperate for relief. A more sensible approach is to qualify the risk with the statement that, notwithstanding the mathematical derivation, the complication of interest has been known to occur. Case reports show spontaneous spinal hematoma and spinal hematoma in patients with normal and abnormal hemostasis function. One approach is to assume that, until proven otherwise, neuraxial anesthesia in patients with thrombocytopenia should carry a spinal hematoma risk that is at least as high as the recognized risk of such complication compiled mainly from patients without hemostatic deficiencies receiving a neuraxial block. In a retrospective chart review of 43,200 epidural catheterizations, 6 confirmed cases of epidural spinal hematoma were identified, yielding an incidence of 0.00014 or 0.014% (95% CI, 0–0.03).13 This 95% CI should also be presented to the patient. By the same argument, although the 1-tailed 95% CI risk of aspiration using inhalational induction in infants with pyloric stenosis based on 252 uncomplicated cases is 0%–1.2%, clinicians should also be aware that 6 cases of aspiration among 26,988 emergency pediatric cases (risk of 0.02%; 95% CI, 0.01%–0.05%) have been reported.14 (This risk is approximate because not all aspirations occur during induction.14,15) The decision to undertake an intervention or a certain approach depends on the risk of complications and its benefits as well as the risk-benefit ratios of alternative approaches or interventions. Take induction of anesthesia for pyloric stenosis as an example. As pointed out earlier, many pediatric anesthesiologists currently induce with sevoflurane, as opposed to propofol and succinylcholine, in a rapid sequence (RSI) with CP. Proponents of inhalational induction typically base their argument on the problems that can be associated with CP—interference with inserting the laryngoscope, potential for worsening the laryngeal view, impeding intubation, obstructing bag-mask ventilation (if required), lack of standardization of the downward force, questionable execution, relaxation of the gastroesophageal sphincter, retching, and so on. These proponents also like to state that there has not been any randomized controlled trial performed to prove that CP is beneficial while (ignoring the irony) citing anecdotal reports of aspiration despite CP.4,11,16 The Scrimgeour et al11 report of 252 uncomplicated inhalational inductions for pyloromyotomy appears to support inhalational induction. Indeed, an anesthesiologist who has done a pyloromyotomy that way every single weekday of the year for a year without complications could be forgiven for his/her confidence. However, objective analysis suggests that the risk is, as discussed earlier, not 0, and clinicians who have witnessed pulmonary aspiration may beg to differ. The alternative to inhalational induction is RSI with CP, which probably has an even better safety record. Senior anesthesiologists may remember the old dogma that discourages assisted ventilation during CP and that CP must never be withdrawn until tracheal intubation is confirmed through auscultation.17 The Scrimgeour et al11 data further dispel that dogma. In inducing anesthesia for pyloromyotomy, many clinicians continue to use RSI with CP, but when laryngoscopy and intubation (and, if needed, bag-mask ventilation) appear to be hindered by CP, they would ask the assistant to relax CP to see whether conditions improve. There are situations in which it is debatable whether to proceed with epidural catheterization in a thrombocytopenic parturient. For example, it is midnight when you are asked to see an obese patient with preeclampsia (platelet count 51,000 mm–3) with an edematous airway with nonreassuring fetal heart rate. The Figure shows a decision tree that reflects what your thought process may be in choosing 1 of 2 options: (1) epidural, confirm good placement, and be prepared to supplement block with lidocaine + bicarbonate in case of emergency cesarean delivery; or (2) no epidural, general anesthesia if crash emergency cesarean is required. The decision tree progresses from left to right, starting with option 1 or 2, and showing the possible chance events (labeled above each branch) emanating from them, with the corresponding probabilities (labeled beneath each branch), and ending with various possible outcomes. Table 2 shows the probabilities of the possible outcomes for the decision tree and the utility values of the final outcomes.Table 2.: Probability and Utility Values to Populate the Decision Tree (Figure)Figure.: Decision tree for choosing between epidural placement and no epidural placement in a patient with thrombocytopenia and preeclampsia with a difficult airway and nonreassuring fetal heart rate (FHR). The probability of each chance event is below each corresponding branch. The square node is a decision node. The round nodes are chance nodes. The triangular nodes are terminal (outcome) nodes. Notice the small difference between the probability-weighted composite utility of the 2 decision options, suggesting that small changes in the probabilities and utilities could potentially change the decision.Based on this analysis, you decide to proceed with epidural placement early, limit the amount of local anesthetic to facilitate close neurologic monitoring, and have platelets on hand should excessive bleeding occur during cesarean delivery. You also note that important details have been left out of this decision tree (and that you are not likely to pull out a calculator or fuss over what CI to use), but it does reflect how our brain weighs the different options and potential outcomes. You also notice that, depending on the probability estimates and utility values, the closeness of the probability-weighted composite utility values of the 2 options suggests that the decision could easily have gone the other way. This finding reflects the differing opinions that one may get from different anesthesiologists presented with the same scenario. If n is large enough, then eventually some rare outcomes do occur. The 95% upper limit confidence for rare events that occur 1, 2, 3, or 4 times are 5.5/n, 7.2/n, 8.8/n, and 10.2/n, or simply 5, 7, 9, 10, respectively.20 Alternatively, one can log on to one of those previously mentioned websites for calculating the binomial CI mentioned earlier. Finally, we offer a word of caution on drawing inferences from uncontrolled retrospective case series. In the Scrimgeour et al11 institution, a small percentage of patients with pyloromyotomy were given intravenous induction and some received RSI with no explanation given, and they were excluded from the analysis. In the analyzed obstetric series, little or no information was provided on whether only experienced staff was allowed to attempt epidurals in thrombocytopenic parturients, the number of attempts allowed, whether catheterization was abandoned if there was blood in the Tuohy needle, and whether parturients who could not keep still were declined. Furthermore, missing data are always a possibility in database analysis. For example, is there any possibility that a patient who developed spinal hematoma was transferred to another center that was not part of the MPOG consortium? Ideally, studies using case series should be prospective and have well-defined a priori objectives and protocol, inclusion and exclusion criteria, timeline, and outcomes. In summary, when dealing with rare events in which the plausible complication of interest has had 0 occurrences after n trials/interventions, one can calculate the upper limit of the 1-tailed 95% CI using the Rule of 3, which is 3/n (for n ≥ 30). For rare events that have occurred once or more times, simply log on to 1 of many websites where the limits can be calculated by inputing the number of events and trials. The upper limit of a 1-tailed 95% CI is a conservative estimate worth considering if the complication of interest is serious. If the number of trials is relatively small, then this approach may produce risk estimates that seem conservative. For a less conservative estimate, readers may wish to read the Quigley et al20 derivation that for 0 complications after n trials, a “more realistic” estimate is 0.4/n.18 Whether to use the conservative or so-called more realistic estimate depends on the benefit-harm ratio of the intervention and that of alternative interventions. Finally, clinicians, patients, and stakeholders must take into consideration-related reports that document the complication in question and include such information in the calculus of their decision. DISCLOSURES Name: Anthony M.-H. Ho, MD, FRCPC, FCCP. Contribution: This author helped draft and revise the manuscript and approved the final version for publication. Name: Adrienne K. Ho, MBBS. Contribution: This author helped draft and revise the manuscript and approved the final version for publication. Name: Glenio B. Mizubuti, MD, MSc. Contribution: This author helped revise the manuscript and approved the final version for publication. Name: Peter W. Dion, PhD, MD, FRCPC. Contribution: This author helped revise the manuscript and approved the final version for publication. This manuscript was handled by: Richard C. Prielipp, MD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.278 | 0.591 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.006 | 0.010 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.003 | 0.007 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.006 | 0.006 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".