Essential Policy Intelligence | Conseils indispensables sur les politiques
Bibliographic record
Abstract
The latest round of the Program for International Student Assessment (PISA) shows statistically significant declines in mathematics scores for most Canadian provinces, and in science and reading scores for many provinces. The PISA background research sheds light on which policies – among many hotly debated approaches – probably do improve student performance. Policies that probably do work: • Pre-primary (early childhood) education improves outcomes among 15-year-old students – especially among socially disadvantaged students. • School autonomy improves outcomes – provided school-level academic results are posted publicly. • Paying secondary school teachers well is associated with better outcomes – clearly evident in a Canada/US comparison. • Subsidizing a sizeable minority of students to attend private schools probably helps explain Quebec’s superior mathematics results – in public as well as private schools. Policies that appear not to work: • If the national student/teacher ratio is already below 20 (as is true in Canada), lowering it further is unlikely to improve outcomes. • Increasing instruction time for mathematics is unlikely, by itself, to improve mathematics scores. This E-Brief benefited from rigorous review by both C.D. Howe Institute analysts and external discussants. I thank Marie-Anne Deussing for her detailed review of the manuscript. Colin Busby and James Fleming contributed editorial advice and organized preparation of the E-Brief.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".