Avoiding ecological fallacy: assessing school and teacher effectiveness using HLM and TIMSS data from British Columbia and Ontario
Bibliographic record
Abstract
There are two serious methodological problems in the research literature on school effectiveness, the ecological problem in the analysis of aggregate data and the problem of not controlling for important confounding variables. This dissertation corrects these errors by using multilevel modeling procedures, specifically Hierarchical Linear Modeling (HLM), and the Canadian Trends in International Mathematics and Science Study (TIMSS) 2007 data, to evaluate the effect of school variables on the students’ academic achievement when a number of theoretically-relevant student variables have been controlled. In this study, I demonstrate that an aggregate analysis gives the most biased results of the schools’ impact on the students’ academic achievement. I also show that a disaggretate analysis gives better results, but HLM gives the most accurate estimates using this nested data set. Using HLM, I show that the physical resources of schools, which have been evaluated by school principals and classroom teachers, actually have no positive impact on the students’ academic achievement. The results imply that the physical resources are important, but an excessive improvement in the physical conditions of schools is unlikely to improve the students’ achievement. Most of the findings in this study are consistent with the best research literature. I conclude the dissertation by suggesting that aggregate analysis should not be used to infer relationships for individual students. Rather, multilevel analysis should be used whenever possible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".