Accounting for heterogeneity in the dependence mechanism of longitudinal data
Bibliographic record
Abstract
Longitudinal data occur frequently in practice where measurements are collected from subjects over time with an aim to understand the dependence mechanisms among these measurements. A major challenge in longitudinal data analysis is the presence of a complex dependence structure due to both between and within individual heterogeneity. This thesis develops new statistical methodologies that incorporate potential heterogeneity in the dependence structure in various longitudinal data problems. In the first part, we introduce a D-vine copula-based heterogeneous dependence model which provides a flexible representation of time-heterogeneous dependence in univariate longitudinal data with a continuous outcome. The proposed model allows for time adjustment in the dependence structure of unequally spaced and potentially unbalanced longitudinal data. We show that the proposed approach offers flexibility over its time-homogeneous counterparts as well as allows for parsimonious model specifications at the tree or vine level for a given D-vine structure. The performances of the time-heterogeneous D-vine copula models are evaluated through simulation studies and by real data from the Manitoba Follow-up Study. In the second part, we propose an approach to incorporate potential heterogeneity in the random effects covariance matrix in longitudinal data with missing responses and mismeasured covariates. The proposed approach uses a modified Cholesky decomposition and allows the random effects covariance matrix to depend on covariates. This decomposition provides an unconstrained and statistically meaningful reparameterization of the random effect covariance matrix which can be modeled without the concern of positive definiteness of the resulting estimators. The performance of the proposed approach is evaluated through simulation studies and is demonstrated using longitudinal data from Framingham Heart Study. In the last part, we review two major statistical models for longitudinal functional data that are spatially correlated and propose a computationally efficient modeling approach by incorporating a spatio-temporal dependence structure in the error process. Numerical experiments are conducted to compare these models and to investigate the impact of ignoring spatial correlation on prediction performance. We discuss the limitations of these models and outline future directions to develop flexible models that can incorporate potential heterogeneity in the dependence structure of spatial longitudinal data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.077 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".