Measuring the Effectiveness of Student Aid, 2005-2008 [Canada]
Bibliographic record
Abstract
The Measuring the Effectiveness of Student Aid (MESA) dataset comprises a sample of low income students receiving student financial aid in 2006-07. Students were contacted first (Cycle I) in February-May of that academic year (the precise date varying by province), and were then followed up in 2007-08 (Cycle II), contacted in February-April of that year. Students will be contacted again in 2008-09 for the last time. The dataset represents a national sample, including all provinces-except for Prince Edward Island. In the spring and summer of 2005, the Canada Millennium Scholarship Foundation negotiated a series of agreements with provincial governments to deliver a set of bursaries (known as “Access Bursaries”) to first-time, first-year undergraduates from low-income families. These agreements are all broadly similar though eligibility criteria vary slightly by jurisdiction (section 1, below, describes the Access Bursaries as they exist in each province). Students do not need to apply for the award separately; instead, they are automatically considered for the award through their application for provincial student assistance. The sample represents a particular subset of the students who received student financial aid in their first year of post-secondary education in 2006-07. In the majority of the provinces, this subset consists of the students who received a Low Income Bursary from the Millennium Scholarship Foundation. In British Columbia and Nova Scotia, a control group made up of students who received financial aid but not the Millennium Bursary was surveyed as well. The Ontario sample is made up those Millennium Bursary recipients who also received a Canada Access Grant and those who did not, with sub-samples selected from each group (all appear together in the data but can be separately identified). The Bursary and the Grant are awarded in similar amounts, but the eligibility requirements are different. This dataset was freely received from the Canada Millennium Scholarship Foundation. Some work was required for the variable and value labels, and missing values. They were corrected as best as possible with the documentation received. Caution should be used with this dataset as some variables are lacking information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.012 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".