Statistics Canada’s Learning Resources: A Key Channel for Educators
Bibliographic record
Abstract
Statistics Canada is the federal agency responsible for collecting information on all aspects of the Canadian society and economy. We publish and make this information available in a variety of formats, including online on a website that boasts more than 1 million visits monthly. More than 40 % of these visits are from educators and students. A special area of our website called Learning Resources is dedicated to providing the education community with theme driven data and articles, hundreds of curriculum based learning activities and expert advice on statistical skills. Through this Learning Resources website and other grassroots initiatives, Statistics Canada is building a relationship with educators to encourage the application of data and data concepts in classrooms across the country. Statistics Canada's role in enabling educators At Statistics Canada, our business is data. Close to 6,000 employees work at perfecting the processes and outputs involved in surveys. Beyond that, Statistics Canada strives to make its data easily understandable to the Canadian people so that they can effectively apply them and make decisions based on them. We have a vested interest in creating an appetite for our data and in making them understandable and easily used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.039 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.009 |
| Science and technology studies | 0.014 | 0.004 |
| Scholarly communication | 0.014 | 0.009 |
| Open science | 0.003 | 0.010 |
| Research integrity | 0.007 | 0.009 |
| Insufficient payload (model declined to judge) | 0.192 | 0.081 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".