STEVE MARKS, Module 8: Becoming a Trusted Digital Repository
Bibliographic record
Abstract
Steve Marks has accomplished something that very few people in the world have: he created a Trusted Digital Repository (TDR) that met the criteria of the Trustworthy Repositories Audit & Certification (TRAC).It was the first repository in Canada, and one of only six in the world.Because of the significance of this task, this publication is important to consider: the author went beyond theorizing how a TDR could be created and actually achieved it.Marks undertook this task when he was the digital preservation librarian at the Toronto-based Scholars Portal, a service of the Ontario Council of University Libraries (OCUL): the Scholars Portal e-journals database TDR passed the very stringent Centre for Research Libraries (CRL) audit and obtained the rare certification in February 2013. 1 To pass the audit and be granted certification, a TDR must demonstrate compliance with the TRAC criteria and the strict "gold standard" of ISO 16363, Audit and Certification of Trustworthy Digital Repositories.Marks defines this ISO as "an internationally recognized set of criteria that can be used to measure the credibility of repositories' specific preservation programs and services" (p.2).The book is published by the Society of American Archivists (SAA) and is part of its Trends in Archives Practice series.I applaud SAA for creating this series: the books are well priced, short (around 100 pages), and available in print, EPUB, and PDF formats.Marks contributed this publication to the series in 2015 in order to share with the archival community his knowledge of TDRs and his experience with audits.The book starts with a note written by editor Michael Shallcross, who provides a short, helpful explanation of why ISO 16363 is important for archives.The introduction by Bruce Ambacher focuses on the history of trustworthiness and the development of the ISO standard.While well written and interesting, Ambacher's chapter might be too detailed for some readers, 1 Marks has since moved on to become the digital preservation librarian at the University of Toronto Libraries, Information Technology Services.larger philosophical, technological, and ethical issues and opportunities, both enchanting and disturbing, that are facing archives in the present and (frighteningly near) future.For these reasons, the text is not only essential reading for those interested in the intersection of art and archives, but is also a rich site for reflection on the nature and capacities of archives in contemporary society.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.047 | 0.029 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".