MétaCan
Menu
Back to cohort
Record W7143863137 · doi:10.24619/00000788

How could archival metadata be successfully shared?: an international survey of online union catalogs

2008· article· ja· W7143863137 on OpenAlexaboutno aff
タカヒロ サカグチ, Takahiro SAKAGUCHI

Bibliographic record

VenueInstitutional Repositories DataBase (IRDB) · 2008
Typearticle
Languageja
FieldComputer Science
TopicLibrary Science and Information Systems
Canadian institutionsnot available
Fundersnot available
KeywordsMetadataUnion catalogDatabase catalogMetadata repositoryMeta Data ServicesGeospatial metadataData sharingNational archives

Abstract

fetched live from OpenAlex

アーカイブズ情報の共有化は、アーカイブズの保存と利用をめぐる一般の理解を広げるのに大きな効果があるという点で重要であり、全国のアーカイブズ機関による協同の成功を期すにあたっては、その方法論を綿密に検討する必要がある。そこで、イギリス、アメリカ、カナダ、オーストラリア、日本において構築されている全国的アーカイブズ情報共有化データベースを対象に、共有化の仕組みとメタデータ標準類の活用方法について調査を行った。調査項目は「記述作成機関」「記述レベル」「記述項目」「検索項目」「記述規則」「シソーラス」「EADの活用」である。その結果、次のような点が明らかになった。1)データの記述は基本的にアーカイブズ資料を所蔵している各所蔵機関の役割とされていた。2)記述のレベルはフォンドに相当する単位を基本としていた。3)ISAD(G)で記述が必須とされている項目に加え、いくつかの項目が記述項目とされていることが多かった。4)作成者名称は、すべての国で検索項目となっていた。5)記述規則、シソーラス、EADの活用といった発展的な課題については、データベースによる相違が大きかった。 Sharing of archival metadata is crucial in order to expand people's understanding of Management of and access to archival resources. Because the collaboration project between various archival institutions would be difficult, it is necessary to consider carefully about its strategy and methodology to develop archival union database successfully. This article is on an international survey of 5 national online union catalogs: A2A(UK), NUCMC(USA), Archives Canada, RAAM (Australia), and National Archival Information Database and Network (Japan). This survey focused on the system for collaboration and the use of metadata standards, especially on who create archival descriptions, levels of description contained, metadata elements, elements used for search, content standards, thesaurus,and use of EAD. The main findings are: 1)Basically, data creation is the role of each institution with archival holdings: 2) All databases contain fonds level description: 3) In addition to the elements identified as mandatory in ISAD(G), some elements are used. 4) Name of creator (s) is used as search elements in all databases. 5) There are wide variations on the use of content standards, thesaurus and EAD.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.013
metaresearch head score (Gemma)0.047
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.988
Threshold uncertainty score0.068

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0130.047
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0130.032
Science and technology studies0.0040.003
Scholarly communication0.0120.016
Open science0.0010.004
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0080.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.083
GPT teacher head0.292
Teacher spread0.210 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2008
Admission routes1
Has abstractyes

Explore more

Same venueInstitutional Repositories DataBase (IRDB)Same topicLibrary Science and Information SystemsFrench-language works237,207