T3010 Canadian Charity Returns Legacy Extract Dataset [2003-2017]
Bibliographic record
Abstract
This dataset is a legacy version of T3010 Canadian Charity Returns Dataset containing year-by-year extracts of directors and trustees, donees, financials, programs, schedules, and section A_C from 2003-2017 as csv and stata files. This dataset might be more useful for researchers wanting to work with specific timespans or using more manual analysis. For a more complete dataset including documentation, syntax, and data from 2003-2022, go to Lars Nielsen; Secure Empirical Analysis Lab (SEAL), 2020, "T3010 Canadian Charity Returns Master Dataset [2003-2022]", https://doi.org/10.5683/SP2/QXWUAZ, Borealis V4. -------------------------------- T3010 Canadian Registered Charity Information Returns data includes information about charities and public/private foundations, financial information, trustees and financial transactions between non-profits. The T3010 is the tax return that must be filed by all registered Canadian charities under the Income Tax Act. Canadian charities are actually tax-exempt, but the CRA collects T3010 data and renders it publicly available so that governments, private donors, and the general public know what charities are doing and how charities are using their resources. Original and more recent data can be requested from the Canada Revenue Agency (CRA) website. SEAL's former Lab Manager Lars Nielsen processed and cleaned the datasets included in this Dataverse. On 2025-06-23 at the transition of SEAL into its next phase, SEAL Lab Manager Lily Wang transferred responsibility for this dataset to RDM Services. This included a formal email, the current Terms of Use Agreement users have signed to now, as well as the contract between Cardus Institute and McMaster. Cardus is a nonpartisan thinktank that contracted SEAL to clean the T3010 data. The information in this package contains public data submitted to the Canada Revenue Agency (CRA) by the respective registered charity in its Form T3010, Registered Charity Information Return. Please note that this data may have been manually entered by the CRA, and it has not necessarily been verified for accuracy or completeness by the CRA’s Charities Directorate. Statistics and data are produced or compiled by the CRA’s Charities Directorate for the sole purpose of providing the public with direct access to public information about registered charities in Canada. The CRA is not responsible for the use and manipulation by any persons of this information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.010 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.013 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.060 | 0.063 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".