Semiparametric Methods for Survival Data with Clustering, Outcome-Dependent Sampling, Dependent Censoring, and External Time-Dependent Covariate.
Bibliographic record
Abstract
In this dissertation, we focus on the development of semiparametric methods for estimating proportional hazards models in the presence of non-standard data structures, namely clustering, outcome-dependent sampling, dependent censoring and external time-dependent covariate. In the first chapter, we propose methods based on estimating equations for case-cohort designs with clustered failure time data. We assume a marginal hazards model with a common baseline hazard and common regression coefficients across all clusters. Compared to their closest competitors in the literature, the proposed methods feature more tractable asymptotic derivations, variance estimation with reduced computational burden, and potentially increased efficiency. We apply these methods to the study of mortality among Canadian dialysis patients. In the second chapter, we propose methods for dealing with failure time data in the setting where the probability of sampling subjects depends on the outcome (e.g., death, survival) and where subjects are censored in a manner which is dependent on the failure rate. We employ a novel double-inverse-weighting scheme which combines weights arising from the probability of remaining uncensored and from the probability of being sampled. The proposed methods are applied to study the wait-list mortality among patients with end-stage liver disease. The third chapter is motivated by the challenges of fitting complex models to data from the smaller countries participating in the Dialysis Outcomes and Practice Patterns Study (DOPPS). We perform a comprehensive investigation of the association between the day-of-week-specific death rates and the dialysis schedule in the U.S., several European countries and Japan. Three Cox models are considered in which 'day of the week', 'day of dialysis schedule', or 'days since last dialysis' serves as a time-dependent covariate. The models are compared and contrasted, with special attention given to the setting where the sample size is small.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".