MétaCan
Menu
← Back to cohort
Record W3197231752

Program evaluation with multilevel longitudinal data: evidence from simulation study and cluster randomized controlled trial

2021· dissertation· en· W3197231752 on OpenAlexaboutno aff
A. R. Hasan

Bibliographic record

VenueMspace (University of Manitoba) · 2021
Typedissertation
Languageen
FieldDecision Sciences
TopicEvaluation and Performance Assessment
Canadian institutionsnot available
Fundersnot available
KeywordsMultilevel modelCluster (spacecraft)Randomized controlled trialLongitudinal dataComputer sciencePsychologyData miningData scienceMedicineMachine learningInternal medicineOperating system
DOInot available

Abstract

fetched live from OpenAlex

Background: Many intervention programs are implemented with cluster randomized controlled trial (cRCT), i.e., the clusters (e.g., classrooms or schools), not subjects, were randomly assigned into treatment or control groups. The outcome variables are also reported for multiple time points (e.g., pre-and post-intervention). The mixed model is commonly used in analyzing longitudinal data, and most research on program evaluation ignored the non-independence of subjects within-cluster even data are from cluster sampling design. Ignoring the dependency between measurements at different times within-subject has been shown that it can lead to the incorrect estimates of standard error and the type-I-error, but the consequence of ignoring non-independence of subjects within-cluster and/or between measurements within-subject has not been investigated extensively. Objectives: The objectives of this study are, (i) to examine the impact of ignoring the within-cluster correlation and/or within-subject correlations on program evaluation in the cRCT studies; (ii) to evaluate the effect of a mental health prevention program with the cRCT design and investigate factors that moderate the successful intervention. Methods: We implemented both simulation and application to real study to illustrate the impact of ignoring non-independence on effect size estimation in the cRCT. Project 11, a prevention program in Manitoba schools to improve mental health, was used as an illustration example for the empirical study. This study has been implemented with cRCT by randomizing the classrooms into treatment or control groups. Three-time repeated measurements of each student clustered within classroom exhibit a three-level hierarchy of data structure. Based on this data, we simulated three-level data with different magnitudes of intraclass correlation to represent different degrees to which individuals resemble each other relatedness within the cluster. We applied both 2-level (ignoring a level, i.e., either within-class correlation or within-subject correlation) and 3-level (considering both correlation terms) multilevel models to compare the outcome of interest with the true population parameters. The Project 11 data was used as an example to illustrate the consequence of ignoring the higher level of hierarchy on the estimation of intervention and moderation effects. Results: The simulation study shows that ignoring the within-cluster correlation and/or within-subject correlations gives less accurate parameter estimates, and the coverage rate also decreases. ii This impact depends on the sample size and ICC of each level of the multilevel data, and for small sample size, the impact is found severe. The results of the empirical data analysis show that both the random effect and fixed effect parameter estimates along with their standard errors get affected if a level is ignored. The analysis of Project 11 data provides evidence of the positive effect of this cRCT based mental health intervention program. The behavioural difficulties of students significantly decrease over time, and socioeconomic status (SES) has a moderation effect on the program outcome. Although gender does not moderate the effect of the intervention program directly, significant gender difference on the moderation effect of SES is observed. Conclusions: In cRCT based study, it is important to consider the within-cluster correlation and/or within-subject correlations as ignoring these correlations gives incorrect results and, therefore, can lead to different research conclusions. Project 11 program effectively reduces participated students’ behavioural difficulties, and SES significantly moderates the outcome. The study provides guidance for school-based program design and evaluation, and we can learn more about how and for whom interventions work.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.177
metaresearch head score (Gemma)0.460
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.823
Threshold uncertainty score0.934

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1770.460
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0070.013
Bibliometrics0.0030.004
Science and technology studies0.0010.003
Scholarly communication0.0030.004
Open science0.0040.003
Research integrity0.0040.005
Insufficient payload (model declined to judge)0.0080.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.268
GPT teacher head0.461
Teacher spread0.193 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueMspace (University of Manitoba)→Same topicEvaluation and Performance Assessment→French-language works237,207→