Multistate Models for Biomarker Processes
Notice bibliographique
Résumé
Multistate models are widely used for describing life history processes. In studies where \nindividuals are observed continuously, the transition times between states are known exactly. However, when individuals are observed intermittently, transition times and even the states visited between successive observations, may be unknown. Irregular intermittent observation is a special case of intermittent observation where the observation times vary across individuals. \n \nIn the case of intermittent observation, we may not be able to estimate model parameters precisely. In the first part of the thesis, we review methods of estimation for \nMarkov models in this situation, and provide a numerical study that shows the loss of \nefficiency in estimation for intermittent observation compared to continuous observation in both progressive and bi-directional multistate models. Then, application to data from the CANOC, Canadian Observational Cohort study of HIV-positive individuals whose virus has been suppressed by combination antiretroviral therapy, illustrates the effect of gap times on estimation efficiency. \n \nIrregular observation is very common in longitudinal data on disease history of individuals in observational studies. However, there are considerable challenges in checking models with these observation schemes, since there is a strong possibility that this irregularity may be induced by the dependency of inter-visit times on previous process history. As a result, followup visits from this kind of data are subject to disease state-dependency, which needs to be taken into account to prevent biased analysis. The second part of this thesis begins with a review on the estimation of marginal process features such as failure time distributions and prevalence probabilities in the context of Markov multistate models with intermittent observations. A method for estimation of these features is developed using Inverse Intensity Weights (IIW). This method corrects the estimation bias due to dependent observation times. Simulation studies illustrate that the proposed method yields estimates that are close to the true values, while the method that ignores the dependency yields estimates that differ substantially from the true values. Then, an application involving viral load dynamics in a group of individuals from the CANOC study is presented. \n \nIn practice, we may want to consider models for which transition intensities depend on \ninternal covariates related to previous process history. There are, however, challenges in fitting and checking models involving internal covariates, and in making predictions. In the third part of this thesis, we have developed an algorithm that simulates possible sample paths of individuals' processes, and we use it for prediction and model checking. \n \nFinally, there has been recent discussion of model assessment of multistate models. \nThere remain, however, some difficulties in model assessment with irregular intermittent \nobservations. The last part of this thesis addresses problems that arise with methods based on comparison of empirical and model-based estimates. We propose the use of likelihood ratio tests within the Markov process family, and methods of estimating the power of these tests are given. We also propose a method for comparing models based on different outcome spaces in terms of prediction. Finally, the proposed methods are applied to a group of individuals in the CANOC study.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,018 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,003 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,004 | 0,004 |
| Science ouverte | 0,003 | 0,003 |
| Intégrité de la recherche | 0,004 | 0,005 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,017 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».