Uncovering Research Potential of Administrative Data on Charitable Foundations in Canada
Bibliographic record
Abstract
This is the first study of its kind to assess the untapped research capacity of administrative data on Canadian foundations. More than twenty years of records collected by the Canada Revenue Agency (CRA) for the entire population of foundations is publicly accessible to researchers. Canadian data offers greater opportunity for nuanced analysis of the charitable foundation sector than information from the comparatively small sample available for U.S. foundations. Despite the richness of Canadian data and the potential it has to inform grantmaking and administrative practices of foundations, the academic community has paid little attention to this wealth of statistical information. This article explores some of the questions that this data can potentially answer. Consultations with foundation representatives help illuminate the directions that the foundation sector would like researchers to pursue with this data.Ceci est la première étude de son genre à évaluer comment certaines données administratives pourraient contribuer à la recherche sur les fondations caritatives canadiennes. En effet, plus de vingt ans de données accumulées par l’Agence du revenu du Canada pour la population entière des fondations sont maintenant accessibles aux chercheurs. Ces données canadiennes représentent une occasion exceptionnelle pour effectuer une analyse nuancée du secteur des fondations caritatives, occasion qui est meilleure qu’aux États-Unis, où l’échantillon est relativement petit. Malgré la richesse des données canadiennes et leur potentiel d’améliorer l’octroi de bourses et l’administration des fondations canadiennes, la communauté académique a porté peu d’attention jusqu’à présent à cette manne de statistiques. Cet article-ci en revanche explore quelques-unes des questions auxquelles ces données pourraient porter des réponses. En outre, des consultations faites auprès des représentants de certaines fondations aident à signaler les directions que le secteur pourrait prendre grâce à ces données.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.214 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.017 | 0.056 |
| Science and technology studies | 0.012 | 0.005 |
| Scholarly communication | 0.012 | 0.003 |
| Open science | 0.003 | 0.009 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".