Multistate Models for Biomarker Processes
Bibliographic record
Abstract
Multistate models are widely used for describing life history processes. In studies where \nindividuals are observed continuously, the transition times between states are known exactly. However, when individuals are observed intermittently, transition times and even the states visited between successive observations, may be unknown. Irregular intermittent observation is a special case of intermittent observation where the observation times vary across individuals. \n \nIn the case of intermittent observation, we may not be able to estimate model parameters precisely. In the first part of the thesis, we review methods of estimation for \nMarkov models in this situation, and provide a numerical study that shows the loss of \nefficiency in estimation for intermittent observation compared to continuous observation in both progressive and bi-directional multistate models. Then, application to data from the CANOC, Canadian Observational Cohort study of HIV-positive individuals whose virus has been suppressed by combination antiretroviral therapy, illustrates the effect of gap times on estimation efficiency. \n \nIrregular observation is very common in longitudinal data on disease history of individuals in observational studies. However, there are considerable challenges in checking models with these observation schemes, since there is a strong possibility that this irregularity may be induced by the dependency of inter-visit times on previous process history. As a result, followup visits from this kind of data are subject to disease state-dependency, which needs to be taken into account to prevent biased analysis. The second part of this thesis begins with a review on the estimation of marginal process features such as failure time distributions and prevalence probabilities in the context of Markov multistate models with intermittent observations. A method for estimation of these features is developed using Inverse Intensity Weights (IIW). This method corrects the estimation bias due to dependent observation times. Simulation studies illustrate that the proposed method yields estimates that are close to the true values, while the method that ignores the dependency yields estimates that differ substantially from the true values. Then, an application involving viral load dynamics in a group of individuals from the CANOC study is presented. \n \nIn practice, we may want to consider models for which transition intensities depend on \ninternal covariates related to previous process history. There are, however, challenges in fitting and checking models involving internal covariates, and in making predictions. In the third part of this thesis, we have developed an algorithm that simulates possible sample paths of individuals' processes, and we use it for prediction and model checking. \n \nFinally, there has been recent discussion of model assessment of multistate models. \nThere remain, however, some difficulties in model assessment with irregular intermittent \nobservations. The last part of this thesis addresses problems that arise with methods based on comparison of empirical and model-based estimates. We propose the use of likelihood ratio tests within the Markov process family, and methods of estimating the power of these tests are given. We also propose a method for comparing models based on different outcome spaces in terms of prediction. Finally, the proposed methods are applied to a group of individuals in the CANOC study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.018 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.017 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".