Characteristics of archival data content standards in North America: a comparative analysis with library cataloguing rules
Bibliographic record
Abstract
諸外国では1980年代以降、アーカイブズに関する情報をいかに記述すべきかを定めた「記述規則」の標準化が進んだ。なかでも北米のものは、先行して制定されていた図書館界の目録規則をベースにしており、両者の比較分析によって、アーカイブズの特性やアーカイブズ学の理論・原則をどのように記述規則に反映させようとしたのかが明確になる。本稿では、北米のアーカイブズ記述規則として米国のDescribing Archives: A Content Standard (DACS)とカナダのRules for Archival Description (RAD)2008年改訂版を、図書館界の目録規則として英米目録規則第2版(AACR2)をそれぞれ取り上げた。各規則の全体的な構成と記述項目の構成について比較した後、「タイトル」「日付」「数量」の項目に関する個々の規定内容について比較分析を行った。その結果、1)「原則の声明」を収録している、2)ISAD(G)第2版が示す記述項目にほぼ対応している、3)タイトルについては記述担当者による補記が前提となっている、4)日付については年・月・日の記載が基本となっている、といった特性が2つのアーカイブズ記述規則に共通してみられた。これらは、記述データの生成をめぐってアーキビストに求められる主体性と密接に関わる点であると思われる。 From the 1980s, archival data content standards were developed in many countries. Standards in North America were produced based on the cataloguing rules ill library community comparing both rules will clearly demonstrate how the characteristics of and principles on archival description have implemented in those standards. The following standards are examined in this article: Describing Archives: A Content Standard (DACS) in USA; Rules for Archival Description (RAD) 2008 version in Canada; and Anglo-American Cataloguing Rules 2nd edition (AACR2) as the library standard. The article analyzes overall structure and descriptive elements of those standards and comparing them with regard to some specific rules on Title, Date and Extent elements. I’ll common characteristics of two archival content standards are: 1) both contain the Statement of Principles; 2) both of descriptive elements are basically compliant with those of ISAD(G); 3) both are based on the assumption that titles are supplied by archivists; 4) dates are entered as not only year but also month and day Such characteristics seem to be closely related to the ability to make decisions which is necessary for archivists creating descriptive data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Scholarly communication Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
| gpt | Scholarly communication Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.022 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".