Five Hundred 5.25-Inch Discs and One (Finicky) Machine: A Report on a Legacy E-Records Pilot Project at the Archives of Ontario
Bibliographic record
Abstract
Il existe un vaste corpus de textes théoriques portant sur le sujet des documents numériques, de leur authenticité et des problèmes liés à leur gestion.Cependant, il existe une lacune au niveau des compte rendus du traitement des documents numériques, surtout en ce qui a trait aux documents numériques patrimoniaux sur disquettes.En 2008, les Archives de l'Ontario ont lancé un projet-pilote sur les documents numériques afin d'analyser les problèmes engendrés par ces disquettes.Jusqu'à maintenant, on a évalué plus de cinq cents disquettes 5,25 pouces, six cents disquettes 3,5 pouces, cinquante cédéroms, deux lecteurs ZIP et plusieurs autres médias.Bien qu'on ait connu un succès considérable avec l'utilisation des lecteurs de disquettes 3,5 pouces externes, le défi de maintenir un lecteur de disquettes 5,25 pouces fonctionnel se poursuit.Ce texte rend compte de : 1) le progrès du projet des documents numériques; 2) les problèmes reliés à la lecture et à l'évaluation des fichiers électroniques, y compris ceux qui ne sont plus supportés ou qui sont de format obsolète; 3) l'investigation légale poussée des logiciels; 4) les nouveaux développements possibles dans le domaine du matériel informatique.L'accent est placé sur la récupération de l'information plutôt que sur la préservation des médias archaïques.ABSTRACT There is a vast amount of theoretical literature on the subject of electronic records, their authenticity, and problems with their management.There is, however, a lack of practical reporting on the processing of electronic materials, especially legacy e-records on floppy disk.In 2008 the Archives of Ontario initiated an e-records pilot project to analyze the issues arising from these floppy disks.To date, over five hundred 5.25-inch floppy disks, six hundred 3.5-inch disks, fifty CD-ROMs, two zip-drives, and a variety of other media have been appraised.Although there has been considerable success in the use of 3.5-inch external disk drives, the struggle to maintain a usable 5.25-inch disk reader has been ongoing.This paper reports on: 1) the progress of the erecords project; 2) the issues involved in reading and assessing computer files, including those of currently unsupported and obsolete formats; 3) advanced software forensics; and 4) possible future hardware developments.The focus is on the recovery of informa tion as opposed to the preservation of archaic media.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.031 | 0.021 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".