ENCODED ARCHIVAL DESCRIPTION WORKING GROUP, SOCIETY OF AMERICAN ARCHIVISTS, and NETWORK DEVELOPMENT AND MARC STANDARDS OFFICE, LIBRARY OF CONGRESS, Encoded Archival Description Tag Library. Version 2002
Bibliographic record
Abstract
Type Definition) developed specifically for the purpose of encoding multilevel archival descriptions using a non-proprietary standard such as SGML (Standard Generalized Markup Language) or XML (Extensible Markup Language).The authors are proud to stress the international effort reflected in this version.The goal of the working group is to keep EAD compatible with ISAD(G) (General International Standard Archival Description); therefore, the 2002 version includes new elements and attributes to accommodate this, and certain existing elements and attributes have been clarified.In addition, according to the authors, the French and German experience with EAD has also been reflected with changes to the EAD structure.This is by definition a technical text, given that it is a reference manual that describes all the tags and elements defined for the EAD DTD, their purpose and intended use, and the possible encoding attributes for each element.The conventions used in the text are explained with the aid of an illustrative figure.Several pages are devoted to a brief review of the types of attributes that are used within the EAD DTD and how such attributes are defined before the text launches into the list of elements.The bulk of the text, 237 pages, is used to define and describe each of the 146 EAD elements in detail.Elements are listed alphabetically by the mnemonic tag names, for example for Acquisition Information is preceded by and followed by .This could be a bit
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.007 | 0.011 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.010 | 0.008 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.156 | 0.079 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".