Approximate Methods For Analyzing Semi-Parametric Longitudinal Models With Non-Ignorable Missing Responses
Notice bibliographique
Résumé
In this thesis, we suggest and explore semi-parametric generalized partially linear mixed models for longitudinal data with non-ignorable and non-monotone missing responses.The key subject of our attention is the estimation of mean response parameters and variance components using a semi-parametric Monte Carlo EM method, where the conditional mean response is semiparametric.We first discuss the penalized regression spline method, which is often referred to as P-splines, for linear mixed model.We investigate the methods for estimating mean response parameters and variance components with complete data.We also investigate the connection between P-splines and linear mixed model through incorporating the non-parametric mean functions into longitudinal linear mixed model.An extensive simulation study using different semiparametric mean response functions are presented.Our simulation study reports that when the true underlying model is partially linear, the penalized spline method provides unbiased and efficient estimators.On the other hand, when the mean response is a correctly specified linear model, the P-spline still provides reliable estimates of the model parameters.Next, we present semi-parametric generalized partially linear mixed models for longitudinal data with non-ignorable missing responses.In this situation, we introduce a parametric model for non-ignorable missing data and incorporate it into the likelihood function.We obtain the asymptotic variances of the proposed estimators by the method of Louis (cf.[2], [7]).In addition, we propose and explore a semi-parametric Monte Carlo EM (MCEM) algorithm for simultaneous estimation of the regression parameters and variance components in partially and in-laws (Ibrahim, AbduAllah, Hesham, Saleh, Yousf, AbduAlaziz, Khawlah and Norah) for their unbelievable love and supports no matter what.For most of all, I would like to thank my loving, supportive, encouraging, and patient husband Waleed Alrajhi whose true support during my Ph.D. studies is very appreciated.The past years have not been an easy journey for both of us in many aspects.I deeply thank him for standing by my side, always and at any cost.Last but not least, my three amazing children, Ali, Mohammed and Almas, who have been an everlasting source of love and enthusiasm.v List of Tables 2.1 Empirical biases, mean squared errors (MSEs), coverage probabilities (CPs), and average lengths of 95% confidence intervals for m = 100 clusters and n = 4 measurements, S = 500. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .23 2.2 Empirical biases, mean squared errors (MSEs), coverage probabilities (CPs), and average lengths of 95% confidence intervals for m = 200 clusters and n = 4 measurements, S = 500. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .23 2.3 Empirical biases, mean squared errors (MSEs), coverage probabilities (CPs), and average lengths of 95% confidence intervals for m = 300 clusters and n = 4 measurements, S = 500. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 2.4 Summary statistics of smoking status for youth group 12 -19 years old. . . . . .29 2.5 ML estimates and standard errors of parameters for CCHS. . . . . . . . . . . . .30 3.1 Empirical biases, mean squared errors (MSEs), coverage probabilities (CPs), and average lengths of 95% confidence intervals of maximum likelihood estimators (MLEs) of regression parameters and variance components, under different sample sizes m = 100, 200 and 300 and n = 2.The simulation is under NMAR model with roughly 23% missing response values, the missing data parameters φ φ φ = (-2.5,0.2, 0.3, 1) t , and m i = M = 2000 Gibbs sampling size, under S = 1000 simulation size. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .64 xii 3.2 Empirical biases, mean squared errors (MSEs), coverage probabilities (CPs), and average lengths of 95% confidence intervals of maximum likelihood estimators (MLEs) of regression parameters and variance components, under different sample sizes m = 100, 200 and 300, and n = 2.The simulation is under NMAR model with roughly 30% missing response values, the missing data parameters φ φ φ = (-3, 0.2, 0.3, 1) t , and m i = M = 2000 Gibbs sampling size, under S = 1000 simulation size. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .65 4.1 Comparison of Henderson's method (H) for mixed P-spline model with EM method for LMM.The simulated biases and mean squared errors (MSEs) are for the estimators of mean responses at five different values of x, and variance components, under 500 simulation size.True mean response m 0 (x) is linear. . .78 4.2 Comparison of Henderson's method (H) for mixed P-spline model with EM method for LMM.The simulated biases and mean squared errors (MSEs) are for the estimators of mean responses at five different values of x, and variance components, under 500 simulation size.True mean response m 0 (x) is quadratic.79 4.3 Comparison of Henderson's method (H) for mixed P-spline model with EM method for LMM.The simulated biases and mean squared errors (MSEs) are for the estimators of mean responses at five different values of x, and variance components, under 500 simulation size.True mean response m 0 (x) is exponential.80 5.1 Comparison of our proposed semi-parametric MCEM method (Method 1) with the MCEM method of Ibrahim et al. (2001) (Method 2).The simulated biases and mean squared errors (MSEs) are for the estimators of mean responses at five different values of x and variance components under 1000 simulation size.Missing data parameters φ φ φ = (-2.5,0.2, 0.3, 1) t leads to NMAR with 23% missing responses.True mean response m 0 (x) is linear. . . . . . . . . . . . . . . . . . .110 xiii 5.2 Comparison of our proposed semi-parametric MCEM method (Method 1) with the MCEM method of Ibrahim et al. (2001) (Method 2).The simulated biases and mean squared errors (MSEs) are for the estimators of mean responses at five different values of x and variance components under 1000 simulation size.Missing data parameters φ φ φ = (-2.5,0.2, 0.3, 1) t leads to NMAR with 23% missing responses.True mean response m 0 (x) is nonlinear (quadratic) . . . . . . . . . .111 5.3 Simulated biases and mean squared errors (MSEs) of the semi-parametric MCEM estimators of mean responses at five different values of x and variance components under 1000 simulation size.Missing data parameters φ φ φ = (-3, 0.2, 0.3, 1) t leads to NMAR with 30% missing responses.True mean response m 0 (x) is nonlinear (quadratic). . . . . . . . . . . . . .
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,024 | 0,102 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,002 | 0,003 |
| Bibliométrie | 0,002 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,003 |
| Communication savante | 0,002 | 0,004 |
| Science ouverte | 0,004 | 0,004 |
| Intégrité de la recherche | 0,002 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».