Equilibrium policy experiments and the evaluation of social programs. Mimeograph
Bibliographic record
Abstract
This paper makes three contributions to the literature on program evaluation. First, we construct a model that is well-suited to conduct equilibrium policy experiments and we illustrate effectiveness of general equilibrium models as tools for the evaluation of social programs. Second, we demonstrate the usefulness of social experiments as tools to evaluate models. In this respect, our paper serves as the equilibrium analogue to LaLonde (1986) and others, where experiments are used as a benchmark against which to assess the performance of non-experimental estimators. Third, we apply our model to the study of the Canadian Self-Sufficiency Project (SSP), an experiment providing generous financial incentives to exit welfare and obtain stable employment. The model incorporates the main features of many unemployment insurance and welfare programs, including eligibility criteria and time-limited benefits, as well as the wage determination process. We first calibrate our model to data on the control group and simulate the experiment within the model. The model matches the welfare-to-work transition of the treatment group, providing support for our model in this context. We then undertake an equilibrium evaluation of the SSP. Our results highlight important feedback effects of the policy change, including displacement of unemployed individuals, lower wages for workers receiving supplement payments and higher wages for those not directly treated by the program. The results also highlight the incentives of individuals to delay exit from welfare in order to qualify for the program. Together, the feedback effects change the cost-benefit conclusions implied by the partial equilibrium experimental evaluation substantially.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.055 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.039 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".