Bibliographic record
Abstract
Database of 7,000 Case FilesThe case files from which the database was created are found in the International Institute of Metropolitan Toronto collection at the Archives of Ontario, Toronto.The finding aid gives the agency's dates as , but it ceased to exist on 31 December 1974 and I found no documents that are dated 1975.A more accurate periodization is 1952-74.The International Institute collection is vast, the materials spanning 18.23 metres.The total number of case files cover the period 1952-72 and make up a significant portion (7.2 metres) of the collection.The case files were generated as follows.When an individual (and on occasion, a couple) came to the Institute seeking assistance with a personal matter (such as finding work, sponsoring the immigration of a child still overseas, or dealing with an elderly mother), a cover form that asked for biographical information and type of "problem" was filled out.(Once established in 1960, the Reception Centre staff carried out this task, though a counsellor or caseworker might also do so.)The subsequent entries on the form and on any additional pages were made by the counsellor during the initial and subsequent appointments with a given client, though a client might meet with more than one counsellor.The secretaries typed up the counsellors' notes, adding pages as they needed, but some counsellors may have sometimes done their own typing.(Some secretaries were also home visitors and probably typed up their own notes.)The database contains 7,000 of these case files for the period 1952-72.The process of selection combined a sampling technique and a specific strategy.Every tenth case file was included for the period 1952-72.In addition, every lengthy file was added in the expectation that they would be more informative.By lengthy, I mean more than three pages of qualitative text, which might consist of a counsellor or caseworker's entries and/or a report(s) or summary report(s) conducted by the social agency or medical or court authorities Appendix
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.033 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.009 | 0.013 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.741 | 0.336 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".