How could archival metadata be successfully shared?: an international survey of online union catalogs
Bibliographic record
Abstract
アーカイブズ情報の共有化は、アーカイブズの保存と利用をめぐる一般の理解を広げるのに大きな効果があるという点で重要であり、全国のアーカイブズ機関による協同の成功を期すにあたっては、その方法論を綿密に検討する必要がある。そこで、イギリス、アメリカ、カナダ、オーストラリア、日本において構築されている全国的アーカイブズ情報共有化データベースを対象に、共有化の仕組みとメタデータ標準類の活用方法について調査を行った。調査項目は「記述作成機関」「記述レベル」「記述項目」「検索項目」「記述規則」「シソーラス」「EADの活用」である。その結果、次のような点が明らかになった。1)データの記述は基本的にアーカイブズ資料を所蔵している各所蔵機関の役割とされていた。2)記述のレベルはフォンドに相当する単位を基本としていた。3)ISAD(G)で記述が必須とされている項目に加え、いくつかの項目が記述項目とされていることが多かった。4)作成者名称は、すべての国で検索項目となっていた。5)記述規則、シソーラス、EADの活用といった発展的な課題については、データベースによる相違が大きかった。 Sharing of archival metadata is crucial in order to expand people's understanding of Management of and access to archival resources. Because the collaboration project between various archival institutions would be difficult, it is necessary to consider carefully about its strategy and methodology to develop archival union database successfully. This article is on an international survey of 5 national online union catalogs: A2A(UK), NUCMC(USA), Archives Canada, RAAM (Australia), and National Archival Information Database and Network (Japan). This survey focused on the system for collaboration and the use of metadata standards, especially on who create archival descriptions, levels of description contained, metadata elements, elements used for search, content standards, thesaurus,and use of EAD. The main findings are: 1)Basically, data creation is the role of each institution with archival holdings: 2) All databases contain fonds level description: 3) In addition to the elements identified as mandatory in ISAD(G), some elements are used. 4) Name of creator (s) is used as search elements in all databases. 5) There are wide variations on the use of content standards, thesaurus and EAD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.047 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.013 | 0.032 |
| Science and technology studies | 0.004 | 0.003 |
| Scholarly communication | 0.012 | 0.016 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".