Structural and formal analysis: the contribution of diplomatics toarchival appraisal in the digital environment
Bibliographic record
Abstract
Introduction ‘Analysis is the essence of archival appraisal’ (Schellenberg, 1956a, 45 [277]). All those who have written about appraisal, regardless of their perspective, beliefs and context, agree that the key to the accurate assessment of the value of records is a systematic and rigorous analysis of their context, interrelationships, form, content and/or use. They may disagree on the methodology for, or on the object of, analysis but, since the mid-19th century, the idea that appraisal could be based on intuition has all but disappeared, being replaced by the conviction that appraisal can only result from a scientific process of analysis, regardless of the interest being served and the criteria being followed. Structural analysis Structural analysis was introduced in the discourse on appraisal in the 20th century by German theorists. Although many archivists in Germany still supported the primacy of content analysis aimed at determining the usefulness of records for future historical research (Zimmerman, 1959), structural analysis began to dominate appraisal methodology, mostly as a consequence of the widespread inter national acceptance of the principle of provenance as the theoretical basis of archival arrangement. If meaning is derived from context, then an understanding of the administrative structure of a records’ creator should be able to guide not simply arrangement, but also appraisal (Heredia Herrera, 1987, 123). To German archivists, the destruction of copies and transitory records was still the proper thing to do, because they were extraneous to the understanding of context and structure (Doehaerd, 1950, 325), until, in 1939, Hans O. Meissner re-issued and developed the systematic appraisal standards formulated in 1901 by Georg Hille. His primary contribution to appraisal methodology was the use of structural analysis to gather an understanding of the organization, functions and activities of the records-creating body. However, he believed that such analysis had to be combined with that of subject content in order to be able to identify records of value (Klumpenhouer, 1988, 52). In 1940, Hermann Meinert endorsed Meissner's standards arguing though that the value of records depends primarily on the significance of a records’ creator within an administrative hierarchy, which can be determined through an analysis of its position in such a structure, of the nature of its activities, and of their relationship with those of superior and subordinate administrative units (Schellenberg, 1956b, 137).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.041 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.004 | 0.053 |
| Scholarly communication | 0.016 | 0.015 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".