Bibliographic record
Abstract
Theses Canada, a service of the National Library of Canada, has coordinated a centralized theses program for Canadian universities since 1965. Our mission has been to acquire and preserve a comprehensive collection of Canadian theses and to provide access to them within Canada and throughout the world. As many Canadian universities move rapidly towards electronic theses submission programs our mission has expanded to include the acquisition and preservation of Canadian theses in this new format and to support Canadian universities during the transition period. To this end a theses portal is being developed for the National Library of Canada website. Our goal is to make the portal a comprehensive repository of freely available Canadian electronic theses and dissertations. Expected to be launched by the end of 2003, phase one of the Theses Canada Online Portal will include bibliographic records for the over 225,000 theses in our collection, as well as approximately 45,000 full text electronic theses digitized for Theses Canada by UMI Dissertations Publishing during the period 1998-2002. Planning for phase two of the portal development, which will permit Canadian universities to submit e-theses and metadata directly to the National Library of Canada, is already underway. The National Library of Canada has been represented on the NDLTD Steering Committee since its inception in 1997. The long term goal of Theses Canada is to make Canadian electronic theses and dissertations available in the NDLTD Union Catalogue.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Scholarly communication Domain: not available · Genre: Other About the Canadian research system: yes · About a Canadian topic: no | Not applicable | low |
| gpt | Scholarly communication Domain: not available · Genre: Other About the Canadian research system: no · About a Canadian topic: yes | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.034 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.012 | 0.033 |
| Science and technology studies | 0.015 | 0.004 |
| Scholarly communication | 0.022 | 0.008 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.503 | 0.312 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".