T3010 Canadian Charity Returns Master Dataset [2003-2022]
Bibliographic record
Abstract
This dataset is a master overview of T3010 Canadian Registered Charity Information Returns from 2003-2022 in CSV and Stata format. It also includes a documentation folder with category code and variable changes over time, as well as a detailed Data Dictionary and PDF of original tax forms. Finally, a complete syntax set of DO files for use in the Stata statistical analysis software is included. Data cleaning for this dataset was conducted in two sets, one for the 2021 release (T3010 2021 Release) and a second completed in May 2024 (2024_T3010_Release). For year-by-year extracts of directors and trustees, donees, financials, programs, schedules from 2003-2017 as csv and stata files, go to Lars Nielsen; Secure Empirical Analysis Lab (SEAL), 2019, "T3010 Canadian Charity Returns Legacy Extract Dataset [2003-2017]", https://doi.org/10.5683/SP2/NSHJFT, Borealis, V3. -------------------------------- T3010 Canadian Registered Charity Information Returns data includes information about charities and public/private foundations, financial information, trustees and financial transactions between non-profits. The T3010 is the tax return that must be filed by all registered Canadian charities under the Income Tax Act. Canadian charities are actually tax-exempt, but the CRA collects T3010 data and renders it publicly available so that governments, private donors, and the general public know what charities are doing and how charities are using their resources. Original and more recent data can be requested from the Canada Revenue Agency (CRA) website. SEAL's former Lab Manager Lars Nielsen processed and cleaned the datasets included in this Dataverse. On 2025-06-23 at the transition of SEAL into its next phase, SEAL Lab Manager Lily Wang transferred responsibility for this dataset to RDM Services. This included a formal email, the current Terms of Use Agreement users have signed to now, as well as the contract between Cardus Institute and McMaster. Cardus is a nonpartisan thinktank that contracted SEAL to clean the T3010 data. The information in this package contains public data submitted to the Canada Revenue Agency (CRA) by the respective registered charity in its Form T3010, Registered Charity Information Return. Please note that this data may have been manually entered by the CRA, and it has not necessarily been verified for accuracy or completeness by the CRA’s Charities Directorate. Statistics and data are produced or compiled by the CRA’s Charities Directorate for the sole purpose of providing the public with direct access to public information about registered charities in Canada. The CRA is not responsible for the use and manipulation by any persons of this information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.008 | 0.020 |
| Science and technology studies | 0.003 | 0.000 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.082 | 0.049 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".