Pretest and james-stein-type shrinkage estimators in cox frailty models
Bibliographic record
Abstract
Survival data analysis may reveal heterogeneity among the observational units in many situations. This heterogeneity is also known as frailty. Analysis that ignores frailty leads to incorrect inferences. The Cox proportional hazards model is commonly used in lifetime analysis to measure the effects of covariates. In some cases, covariates may fail to capture true risk differences in the population. The model might ignore covariates that could explain random frailty, but this can be explained by the existence of covariates not considered in the model. Frailty can prevent underestimation and overestimation of parameters, and it can also be used to estimate the effects of covariates on the response variables. This thesis explores the pretest and shrinkage estimation methods for the Cox proportional hazards frailty model for right-censored clustered survival data when some of the covariates in the model may not be relevant for accurately predicting survival times. In order to address this, we employ two models: an unrestricted model that encompasses all covariates, and a restricted model that includes a smaller set of covariates. For the unrestricted model, we apply the penalized partial likelihood to obtain the unrestricted model estimator. We maximize the penalized partial likelihood under the restriction on the irrelevant covariates to obtain the restricted model estimator. We optimally combine these estimators to obtain pretest and shrinkage estimators. Mean squared error (MSE) and relative MSE are calculated to assess the performance of these estimators. By conducting simulation studies and applying the proposed methods to a lung cancer data, we demonstrate that the shrinkage estimators perform better than the full model estimator when the shrinkage dimension exceeds two. The effects of increasing sample size, number of insignificant covariates, and censoring percentages are also studied through Monte Carlo simulations. The Cox proportional hazards and parametric models are also compared using normal-deviate residuals and pretest and shrinkage methods, focusing on their characteristics from a practitioner's perspective and their practical applications rather than on their theoretical characteristics. This is evaluated using extensive simulation studies and by means of real data sets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".