The Local News Data Hub: Championing data journalism and equity, one story at a time
Bibliographic record
Abstract
The Local News Data Hub at Toronto Metropolitan University supports local journalism at a time when many newsrooms lack the capacity to produce data-informed stories. Once the editorial team identifies data sets that can be used to generate stories for multiple places, student reporters produce a story template that is customized with relevant data for different communities. This allows us to supply newsrooms with free data-driven stories, support/collaborate with journalists/newsrooms working on data projects, and employ/train student journalists. Data Hub stories, which are distributed by The Canadian Press wire service and published on Hub’s website, have used scientific projections for stories on the local impact of climate change and analyzed internet speed-tests to investigate internet service quality in rural areas. In each case, more than 20 news organizations published one or more stories.<br> <br> Our current projects focus on (i) income inequality in Canada and (ii) the country’s aging communities.<br> <br> i) Using data from the Statistics Canada 2021 Census, we looked at income inequality using the Gini coefficient for after-tax income across Canada. The stories highlight the cities/towns in census metropolitan areas that have the greatest income inequality and investigate its consequences.<br> <br> ii) Using Statistics Canada data, we identified 100 census subdivisions with a population greater than 10,000 where at least 25% of the population is 65 or older. Our stories focus on the dozen or so places with high proportions of older people - places such as Parksville, B.C. (45%), Cape Breton, N.S (26%) and Elliot Lake, Ont. (41%) - and ask how prepared they are for the gray tide washing over them.<br> <br> The Data Hub combines data with reporting on human experience to produce stories that point to inequities and advance social justice while also supporting local newsrooms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.006 | 0.000 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.005 | 0.018 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".