Bibliographic record
Abstract
Equivalent File Data files from the 1993-2001 SLID surveys may be analyzed through special arrangement with Statistics Canada through the auspices of the CNEF staff. Please contact CNEF staff by e-mail at CNEF@cornell.edu for further details. Please also note that the years for the Canadian SLID data refer to the reference year in which the data were generated not the year in which the survey was done. This difference is maintained because all SLID documentation here and on the Statistics Canada website use the year data were generated as the reference year. This year we will be releasing a supplementary file of health variables that can be compared across the BHPS, GSOEP, and PSID files. The construction of these variables, funded by the National Institutes on Aging, was not completed in time for our target release date so, for this year only, we will send these variables as supplementary files to users who request them. Please contact CNEF staff by e-mail at CNEF@cornell.edu to request these data. A list of the new health variables will be posted on our web page. The PSID data in this release of the Cross-National Equivalent File 1980-2002 are the same as those included last year because the 2003 data are not yet available. In a departure from past practice, income variables with missing values in the original GSOEP are imputed using a variation of the row and column imputation procedure (Little and Su, 1989). The imputation procedures are
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.005 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.886 | 0.800 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".