Bibliographic record
Abstract
Fifteen years ago, Elizabeth Diamond described the archivist as a forensic scientist. In the past few years, several archival writers have referred to professionals responsible for keeping digital records as trusted keepers or custodians. Undoubtedly, in the digital environment, record professionals are increasingly called to assess and preserve the authenticity of the records they are responsible for, and to act as neutral third parties. But, are they qualified to fulfill this role? This article aims to begin identifying the body of knowledge that a trusted record professional needs in order to assess the trustworthiness of digital records and ensure that their continuing authenticity can be demonstrated, if required, at any point during their life cycle. To do so, it presents some of the concepts developed by the InterPARES Project in the area of diplomatics of digital records; compares them with the relevant concepts of a relatively new discipline called digital forensics; discusses the methodologies used by the two disciplines; and proposes areas that can be jointly investigated by diplomatics and forensics experts to develop an integrated body of knowledge that might be called Digital Records Forensics.RÉSUMÉ Il y a quinze ans, Elizabeth Diamond décrivait l’archiviste comme un scientifique médicolégal. Depuis quelques années, plusieurs auteurs dans le domaine de l’archivistique ont qualifié les professionnels responsables de la préservation des documents numériques de conservateurs de confiance (« trusted keepers »), ou de gardiens (« custodians »). Sans doute, dans l’environnement numérique, on fait de plus en plus appel aux professionnels de l’information pour évaluer et préserver l’authenticité des documents dont ils sont responsables, et pour agir en tant que tierce parties neutres. Mais sont-ils qualifiés pour remplir ce rôle? Cet article tente d’identifier les connaissances que doit avoir le professionnel d’information de confiance pour être capable d’évaluer la véracité (« trustworthiness ») des documents numériques et pour assurer que leur authenticité puisse être démontrée, au besoin, à n’importe quel point dans leur cycle de vie. Pour ce faire, l’article présente des concepts développés par le projet InterPARES dans le domaine de la diplomatique des documents numériques; il compare ceux-ci aux concepts pertinents dérivés d’une discipline relativement nouvelle, le numérique médicolégal (« digital forensics »); il discute des méthodologies dont se servent les deux disciplines; et il propose des domaines qui pourraient être explorés conjointement par les experts en diplomatique et en numérique médicolégal afin de développer un corpus de savoir intégré que l’on pourrait nommer la science médicolégale des documents numériques (« Digital Records Forensics »).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.022 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.006 | 0.017 |
| Scholarly communication | 0.020 | 0.021 |
| Open science | 0.002 | 0.011 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.017 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".