LAURA MILLAR, The Story Behind the Book: Preserving Authors' and Publishers' Archives
Bibliographic record
Abstract
Archivaria, The Journal of the Association of Canadian Archivists -All rights reserved very accessible to amateur archivists, it does offer a strong and compelling message regarding the serious challenges that digital records and the Web pose to both our profession and to amateur family archivists.Since the intent of the Web has been to remain as open and democratic as possible, perhaps the populist approach that Cox promotes is ultimately the best remedy.It is evident that the time has come for archivists to share the responsibility for preserving personal records with the public, which can be accomplished by empowering and better preparing our constituents to contend with the difficulties and perils of preserving their digital treasures on their computers and on the Web.Despite the fact that we may regard this type of memorabilia as overly voluminous and non-archival if it falls outside of our institutional mandates, as Cox persuasively argues, these types of records do hold tremendous value to their custodians and to society.By abandoning our traditionally narrow perspective, we can share our unique skills with a segment of society that may never have stepped inside an archives before, resulting in the broadening of our clientele and support base, and subsequently educating the public about how to take control over their family archives.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.018 | 0.009 |
| Scholarly communication | 0.016 | 0.011 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.011 | 0.013 |
| Insufficient payload (model declined to judge) | 0.012 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".