MétaCan
Menu
Back to cohort
Record W2167268972 · doi:10.1093/ije/dyu105

Avoiding blunders involving 'immortal time'

2014· article· en· W2167268972 on OpenAlexaff
James A. Hanley, Bethany J. Foster

Bibliographic record

VenueInternational Journal of Epidemiology · 2014
Typearticle
Languageen
FieldMathematics
TopicCOVID-19 epidemiological studies
Canadian institutionsMcGill UniversityMcGill University Health CentreMontreal Children's Hospital
Fundersnot available
KeywordsStatistical evidencePsychologyMedicineStatisticsMathematics

Abstract

fetched live from OpenAlex

As Groucho Marx once said ‘Getting older is no problem. You just have to live long enough’. (Queen Elizabeth II, at her 80th birthday celebration in 2006) This award proves one thing: that if you stay in the business long enough and if you can get to be old enough, you get to be new again. (George Burns, on receiving an Oscar, at age 80, in 1996) (Richard Burton died, a nominee 6 times, but sans Oscar, at 59. Burns lived to 100, so how much of the 41 years’ longevity difference should we credit to Burns’ winning the Oscar?) Some time ago, while conducting research on U.S. presidents, I noticed that those who became president at earlier ages tended to die younger. This informal observation led me to scattered sources that provided occasional empirical parallels and some possibilities for the theoretical underpinning of what I have come to call the precocity-longevity hypothesis. Simply stated, the hypothesis is that those who reach career peaks earlier tend to have shorter lives. (Stewart JH McCann. Personality and Social Psychology Bulletin 2001;27:1429–39) Statin use in type 2 diabetes mellitus is associated with a delay in starting insulin. (Yee et al. Diabet Med 2004;21:962–67) For almost two centuries, teachers have warned against errors involving what is now called ‘immortal time.’ Despite the warnings, and many examples of how to proceed correctly, this type of blunder continues to be made in a widening range of investigations. In some instances, the consequences of the error are less serious, but in others the false evidence has been used to support theories for social inequalities; to promote greater use of pharmaceuticals, medical procedures and medical practices; and to minimize occupational hazards. We use a recent example to introduce this error. We then discuss: (i) other names for it, how old it is and who tried to warn against it; (ii) how to recognize it, and why it continues to trap researchers; and (iii) some statistical ways of dealing with denominators measured in units of time rather than in numbers of persons. Patients whose kidney transplants (allografts) have failed must return to long-term dialysis. But should the failed allograft be removed or left in? To learn whether its removal ‘affects survival’, researchers1 used the US Renal Data System to study ‘a large, representative cohort of [10 951] patients returning to dialysis after failed kidney transplant’. Some 1106, i.e. 32% of the 3451 in the allograft nephrectomy group, and 2679, i.e. 36% of the 7500 in the non-nephrectomy group, were identified as having died by the end of follow-up. Patients in the two groups differed in many characteristics: to take into account a ‘possible treatment selection bias’, the authors constructed a propensity score for the likelihood of receiving nephrectomy during the follow-up. They used this together with other potential confounders to perform ‘multivariable extended Cox regression’’. The main finding of these analyses was that ‘receiving an allograft nephrectomy was associated with a 32% lower adjusted relative risk for all-cause death (adjusted hazard ratio 0.68; 95% confidence interval 0.63 to 0.74)’. In their discussion, the researchers suggest that their findings of ‘improved survival’ after allograft nephrectomy ‘challenge the traditional practice of retaining renal allografts after transplant failure’. The title of the article (‘Transplant nephrectomy improves survival following a failed allograft’) suggested causality. They emphasized the large representative sample and the extensive and sophisticated multivariable analyses, but they did caution that ‘as an observational study of clinical practice, their analysis remains susceptible to the effects of residual confounding and treatment selection bias’ and that ‘their results should be viewed in light of these methodologic limitations inherent to registry studies’. They suggested that a randomized trial to evaluate the intervention in an unbiased way would be appropriate. Similar concerns about residual confounding and selection bias, and the need for caution, were expressed in the accompanying editorial reiterating the limitations of the ‘retrospective interrogation of a database’. ‘Residual confounding’ may be a threat, but both authors and editorialists overlooked a key aspect of the analysis, one that substantially distorted the comparison. The overlooked information is to be found in the statements that: 3451 received nephrectomy of the transplanted kidney during follow-up; the median time between return to dialysis [the time zero in the Cox regression] and nephrectomy was 1.66 yr (interquartile range 0.73 to 3.02 yr). (Paragraph 1 of Results section) Overall, the mean follow-up was (only) 2.93 ± 2.26 yr. (Paragraph 3 of Results section) Since the 3451 patients who ultimately underwent a nephrectomy (the ‘nephrectomy group’) had to survive long enough to do so (collectively, approximately 6700 patient-years, based on the reported quartiles of 0.73, 1.66 and 3.02 years), there were, by definition, no deaths in these 6700 pre-nephrectomy patient-years. In modern parlance, these 6700 patient-years were ‘immortal’. There was no corresponding ‘immortality’ requirement for entry into the ‘non-nephrectomy group’. Indeed, all 10 951 patients returning to dialysis after failed kidney transplant began follow-up with their ‘failed graft in place’. Some 7500 of these remained in that initial state until their death (for some, death occurred quite soon, before removal could even be contemplated) or the end of follow-up, whereas the other 3451 spent some of their follow-up time in that initial state and then changed to the ‘failed graft no longer in place’, i.e. post-nephrectomy, state. How big a distortion could the misallocation of these 6700 patient-years produce? The article does not have sufficient information to re-create the analyses exactly. Figures 1 and 2 show a simpler hypothetical dataset which we constructed to match the reported summary statistics quite closely. It was created assuming no variation in mortality rates over years of follow-up or between those lived in the two states. The ‘virtual’ intervention was set up ‘retroactively’ and was limited to the dataset itself, rather than to real individuals, and so could not have affected (other than randomly) the mortality rates in the person-years lived in each state. Excerpts from the simulated mortality experience in the contrasted (‘organ intact’ vs ‘organ removed’) states. Hypothetical lifelines were generated to have an average mortality rate of 3785 deaths in (10 951 × 2.93 = 32 086) patient-years (PY), i.e.,11.8 per 100 PY (as in the actual nephrectomy study1), but with no variation over years of follow-up, and no difference (other than random) between states (‘name intact’ or ‘name removed’). We constructed the dataset by first generating names for 10 951 fictional persons, then distributing the numbers of new cohort entries in a smooth decreasing pattern over the 11 calendar years, and then applying the death rate of 11.8 per 100 PY to the various resulting lengths of available follow-up, until the total number of deaths matched the reported 3785 and the number of PY of follow-up matched the reported 32 086. The 10 951 hypothetical lifelines (3785 completed, 7166 censored) were then ordered from shortest to longest. Finally, starting from the day of return to dialysis and working forward, each follow-up day a number of persons were chosen randomly from among those who had not already been selected, were still alive and were being followed that day. These persons were designated to undergo an electronic ‘removal’ whereby, within the database, just their names (not their failed transplants) were (electronically rather than surgically) removed. The timings of these ‘removals’ (3471 in all) were set so that the median and quartiles of the delay between return to dialysis and becoming nameless matched the delays in the article. The selections, made by a random number generator in 2012, were made blindly, in a retroactive lottery, applied in a forward direction, beginning in January 1994, to lifelines that had already run up to December 2004. Just as in Leibovici16 and in Turnbull et al.,17 these interventions were limited to the 2012 computer file, and could not have affected the comparative mortality rates. Shown are 30 such lifelines selected systematically from these 10 951 ordered hypothetical ones, with a completed lifeline indicated by a single straight line, and a censored one by a pair of lines forming an arrowhead. The timing of the name removal is indicated by an x, and the post-intervention PY by red rather than grey boundary lines. Mortality rates and rate ratios produced by the (A) mis- and (B) proper allocation of pre-‘intervention’ patient years. As explained in Figure 1, the hypothetical data for the 10 951 patients were constructed to have an average mortality rate of 3785 deaths in (10 951 × 2.93 = 32 086) patient-years (PY), i.e. 11.8 deaths per 100PY (as in the actual study), but with no variation over years of follow-up, or between states (no, or pre-‘intervention’ (white background) and post-‘intervention’ (pink background). Indeed, the selection of those who changed states (from white to pink polygon, in B) was made at random, and retroactively. The time location (relative to when the allograft failed) of each death is indicated by a black dot. In B, upper panel, the number being followed at any time is smaller than 3451 because some who had received the ‘intervention’ were already dead before the last ones received it. Figure 2A shows that even though the data were generated to produce the same mortality rate of 11.8 per 100 PY (person-years) in the person-years in the initial and post-‘intervention’ states, the inappropriate type of analysis used in the paper, applied to these hypothetical data, would have resulted in a much lower rate (6.4) in the ‘intervention group and a much higher one (17.1) in the ‘non-intervention’ group. The reason is that none of the 1031 deaths post-‘intervention’ could have occurred, and none of them did occur, in the pre-‘intervention’ PY that are in the to the rate of the 1031 post-‘intervention’ deaths occurred in the post-‘intervention’ the deaths occurred not in but rather in the much of = PY lived in the initial state. The of the PY from the led to the higher than it should have of it was because of these PY they had already that the 3451 patients to have the in other it may not have been that they lived longer because they underwent the but rather that they underwent the ‘intervention’ because they long enough to undergo it. can with some for long of they this bias’ in a suggested that the mortality in could be reported the who die have In this of the data, with no the inappropriate analysis led to an rate ratio of = The corresponding of and an ‘improved survival’ of at years the first 11 years of the of vs years), would have been as having been produced by the whereas they are of the misallocation of the Figure shows an of mortality rates in states. of to the state in which the death would have been should it at that the rates are from random the theoretical rates used to these hypothetical The theoretical rates as over the follow-up years. In the PY in each of follow-up time would be by who were almost one older than the who PY the and so the mortality rates in would be the person-years in the post-‘intervention’ state are a summary rate ratio matched of follow-up time would be than a rate would need to match the person-years on As the between with its on the time and with its on the time does not these but the end of this we a to some we on this against this error at to the when and that: and are by persons in it no of to that mean age at or the age at which the number of deaths be in the of and and it were an into the of the of the on that the mean age at death of and was of of of of of ages still and that the ages of and of of years’ and differed to an or greater a may no be made on of those but and whose mean age at death was It would be almost to them and the of their a the reason for the longevity before they have while may die at any age from their the the of the of to in the of the of and persons in of the time during which the are first in one group and then in the that in the there were on persons and persons. The number of are within these two groups over the calendar and the are This is a so long as the two groups were during the calendar to no or But as in practice, persons are being during the of the the of time at which they the group is into years had the of to when age to age the longevity of with that of the The show that by no a but on the a at age this is it is that all are from age while in actual some do not in a and (as time allocation of follow-up time in the et al. to with into and years. the based on this in and We that the ‘immortal had been used by in the but is the first we of to have the in in examples all the allocation of such with no example of the consequences of The two of the and do have an the difference between and rates based on two groups of and persons, and state the a study has a for a of time before a is to be in the the time during which the is being should be from the of They an to the it is in this that are the ‘immortal in the title of a the Since than a and by and have the number of ‘immortal errors in this cohort in these was at the time of or a medical The were created by the patients into those who were a at some time during follow-up and those who were not in clinical but not all received it at entry to the each follow-up time is into the the of be and the it the of the In of their and use other real to the same and show the consequences of the other examples of by other of in other even with in to 1 before it received the the study of transplant the of the of an how such a blunder can be and 1 some ways to recognize time and to the associated We that some of the from the the to as a were at entry and remained when a authors to the treatment group and the group, rather than to the time when the patients were in the treatment or or states. This may the that many of can be by group in the effects of and use while or use or on in are and their statistical results are to show and in than are those that use to recognize time to recognize time Just as in the of it is that persons in many denominators of time by persons, but time and time is just as are any other or denominators that produce are not Despite many are less with an time into and than they are with research in or than are in the of time used by We forward to the researchers with to their information on the location of so that they can the of and and the rates of in The researchers in an time same with the less the risk of time It is that are than many researchers in the social than in and ratios of are the one is the of that have been completed, it may not much whether one the by their number or the total number by the total (the average number of deaths per of once one to the rate of within just of these the between states, and censored and all it much to stay with the the time The between age and death age for based on age and the to some birthday or in the of those for an are examples of the limitations of with the average that is to to the who rates within the have much than those who to average such as the precocity-longevity hypothesis are and have a But some of this may be a of the of the can the if Groucho Marx were to it, and longevity as the In any (as we in of the longevity for no how or results would be if they were and should be they statistical results that to be to recognize time errors to consequences that in some may be and and not In a the of and the of the can be in the are not the longevity in the of is still available one is to for The does not any does the is an to by Indeed, when one of to the of the of to the that the finding of ‘improved survival’ following allograft nephrectomy was an was that the did not have a to the but that it would on the concerns to the the is of the to get it it and to to time errors from the can ‘immortal errors by into states, rather than persons who ultimately an state into and the that they have been in those groups from the as not Since the Cox is used in this of first of this a analysis with after from a article on survival and the associated computer can be found on the In that we show how we would with these data in a multivariable or with over we found it to with the already in in those in article. We use of to with between is the of what is now as for We the mortality rates and rate ratios This was by the and of and by the of of this we whether we were to the We did not suggested that to such do not in the is than in some of the clinical The article that is the of the in this The in the that led to a but to be hazard ratio of even the was not just by the authors but by their and and and The to how time a in their and to to about their in to the by and We do but after we first some It not be to the in the but we that those in the of this each do their the its to a follow-up that the In of analyses that still have the for the and for their on the longevity an do it just to than the limited of have with the It is one to the a reason to about winning an and years to it is quite to the of the difference in median longevity of years in the left of Figure 1 of the that an even longevity is to the years do not in the but in the In it, one of the authors emphasized that they could not the but that numbers as such do not This what one call a type an set up statistical not is the We on the but results in 3 of the and first the that led to the difference of years and the hazard ratio of How could these analyses have received the and an There was a that the authors on some that such analyses there is a about between the and the The use of for two of the but Cox for the should have to to In the were all so the already in about their may be the that have in from the various and with no of but Cox that the quite pattern of hazard ratios in the lower left of Figure at to those not in the the years and even the adjusted hazard ratio of way to be the mortality rate ratio is approximately the other as with the of of who are or of the of being or of a failed the hypothesis has a to it, and there is other based on the survival of those with et the authors had used a way to study it, and used sophisticated statistical with extensive The in those age less less was as evidence in support of the hypothesis. way to for time is to study an that should have no with the of and to be if the hazard ratios are one may the between a and the of as we do The article has of what was to the results in The statistical matched each one and persons to as but in the results it is to as a the authors do that and following a of or into the analysis, whereas before were together with the all-cause mortality hazard ratios of and suggest they called it, their analysis may at the we the years, and the adjusted and even the hazard ratios in the analyses do The authors now that the analyses in to the end of the Results and not are the one the adjusted from the matched study as the to one could it into for the by that it into a longevity difference of about rather than years. To show how time a in their and to to if the analyses are of it, we the between an and death from any and and the information to we a number of in the by ages years and calendar years to be winning the is the and The dataset available in the Mortality has a total of (the age and of in the or of the on we simulated a that selected some of them to be The of was an with the same as the of in so that the total number of and the average age of were to the of and the average age of in the article. The was that the had to be alive at the time of each other there in large a that other its the could not their just because of this be when we used the same analysis as in Figure 1 in the we a difference in median longevity of years a hazard ratio of with a the of × the hazard ratios we found in the to those in the lower left in the Figure when (as the authors do in their we the age and that who the into the analysis as not having we get to those in the in the to and just two years hazard ratios were not they from at age to at age The reason for the residual is by definition, a who the at age is for years of the age To this one to the to an so is to a in the Cox with risk at the the This is the way to with states rather than the Cox on the authors did who have to the same This may have led them and the to that all was now But with is not one must and each as time and the risk we must one example to the by to the i.e. be in the of or the of to an have it be or a an or a of one to have lived long enough the in to do such longevity requirement is on entry to the

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.251
metaresearch head score (Gemma)0.642
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.749
Threshold uncertainty score0.924

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2510.642
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0040.003
Bibliometrics0.0050.005
Science and technology studies0.0070.078
Scholarly communication0.0110.033
Open science0.0060.013
Research integrity0.0150.029
Insufficient payload (model declined to judge)0.0100.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.315
GPT teacher head0.477
Teacher spread0.162 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations62
Published2014
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of EpidemiologySame topicCOVID-19 epidemiological studiesFrench-language works237,207