MétaCan
Menu
Back to cohort
Record W2245046632 · doi:10.1093/ije/dyv300

Immortal time bias. Response to: Achinger, Go and Ayus

2015· letter· en· W2245046632 on OpenAlexaff
James A. Hanley

Bibliographic record

VenueInternational Journal of Epidemiology · 2015
Typeletter
Languageen
FieldMedicine
TopicEpilepsy research and treatment
Canadian institutionsMcGill University
Fundersnot available
KeywordsMedicinePsychology

Abstract

fetched live from OpenAlex

I concur (as best I can with the indirect evidence) that the authors did carve up time correctly. 1 It is unfortunate that the record had to be corrected this way. It would have been avoided had key details (or even just the one key phrase ‘time-dependent’) not been editorially excised from the manuscript submitted to Journal of American Society of Nephrology ( JASN ). 2 Post-publication, my co-author raised her concern about immortal time with the editor of JASN . She was told that the journal did not have a correspondence section. Her e-mail, which the editor said would be forwarded to the authors, was apparently not received by them. When the authors contacted me and informed me of my incorrect assumptions about their paper, 3 I asked to see some SAS code, but was told that none could be located. Subsequently, Dr Go shared with me the original draft of Table 1 containing the key phrase ‘adjusted hazard ratio comparing time-dependent receipt vs. non-receipt of renal allograft nephrectomy on death from any cause’. This was replaced in the published version by ‘adjusted HR for death for nephrectomy versus non-nephrectomy’. He also told me that the more lengthy description of their modelling strategy was deleted by the journal staff; and he pointed to similar analyses of time-dependent covariates in previous articles he had co-authored. I agreed to contact the IJE staff and tell them that it seemed that the main hazard ratio was indeed based on a proper division of each patient’s follow up time into pre- and (if the allograft was removed) post-removal time. But before doing so, I had two queries for him. One was whether the two denominators behind Figure 2 were numbers of persons or (more appropriately) numbers of person-years: the reported rates were 32 and 36 per 100 person-years, but these did not seem to fit with the reported amounts of follow-up. Since the crude percentages turn out to be 32% and 36%, I wondered if the label in Figure 2 should have read ‘percentage’ rather than deaths ‘per person-year’. The second (also time-related) query was how follow-up time was dealt with in Figure 3, which reported that 10% of those who did and 4.1% of those who did not receive a transplant nephrectomy received a second transplant—a difference that surprised the authors, but for reasons that ‘cannot be determined from our analysis’. My concern was that the durations of follow-up of these two groups differed substantially. If one corrected for this and used the same (time-dependent?) propensities they computed when addressing the primary outcome, and used a time-Cox model with time-dependent covariates, the difference in the adjusted percentages or rates might be even greater. The authors could have used their data to address this second query, and even assess how much of the better survival was mediated by the second transplant. I did not receive a reply to these two queries. These queries bear on the quality and strength and interpretation of the evidence behind an article whose title says a procedure ‘improves’ survival. They deserve to be addressed, and in the subject-matter journal with which these ‘time’ issues were first raised. One lesson from this case is that, although we may not have a lot of control over editorial staff, when we are authors we should insist that key elements are not excised, even if that means removing other less critical material. We do have control over a second aspect: we should keep all computer codes, computer output and data.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.037
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.998
Threshold uncertainty score0.082

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.037
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0020.001
Scholarly communication0.0020.002
Open science0.0010.001
Research integrity0.0260.013
Insufficient payload (model declined to judge)0.0250.013

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.114
GPT teacher head0.425
Teacher spread0.311 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2015
Admission routes1
Has abstractno

Explore more

Same venueInternational Journal of EpidemiologySame topicEpilepsy research and treatmentFrench-language works237,207