MétaCan
Menu
Back to cohort
Record W2525596464 · doi:10.18438/b81s7n

Metadata Quality in Institutional Repositories May be Improved by Addressing Staffing Issues

2016· article· en· W2525596464 on OpenAlexvenueno aff
Elizabeth Stovold

Bibliographic record

VenueEvidence Based Library and Information Practice · 2016
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsnot available
Fundersnot available
KeywordsMetadataStaffingCatalogingQuality (philosophy)Computer scienceMeta Data ServicesDemographicsSample (material)World Wide WebMetadata repositoryMedicineDemographySociologyNursing

Abstract

fetched live from OpenAlex

A Review of: Moulaison Sandy, H., & Dykas, F. (2016). High-quality metadata and repository staffing: Perceptions of United States–based OpenDOAR participants. Cataloging & Classification Quarterly, 54(2), 101-116. http://dx.doi.org/10.1080/01639374.2015.1116480 Objective – To investigate the quality of institutional repository metadata, metadata practices, and identify barriers to quality. Design – Survey questionnaire. Setting – The OpenDOAR online registry of worldwide repositories. Subjects – A random sample of 50 from 358 administrators of institutional repositories in the United States of America listed in the OpenDOAR registry. Methods – The authors surveyed a random sample of administrators of American institutional repositories included in the OpenDOAR registry. The survey was distributed electronically. Recipients were asked to forward the email if they felt someone else was better suited to respond. There were questions about the demographics of the repository, the metadata creation environment, metadata quality, standards and practices, and obstacles to quality. Results were analyzed in Excel, and qualitative responses were coded by two researchers together. Main results – There was a 42% (n=21) response rate to the section on metadata quality, a 40% (n=20) response rate to the metadata creation section, and 40% (n=20) to the section on obstacles to quality. The majority of respondents rated their metadata quality as average (65%, n=13) or above average (30%, n=5). No one rated the quality as high or poor, while 10% (n=2) rated the quality as below average. The survey found that the majority of descriptive metadata was created by professional (84%, n=16) or paraprofessional (53%, n=10) library staff. Professional staff were commonly involved in creating administrative metadata, reviewing the metadata, and selecting standards and documentation. Department heads and advisory committees were also involved in standards and documentation selection. The majority of repositories used locally established standards (61%, n=11). When asked about obstacles to metadata quality, the majority identified time and staff hours (85%, n=17) as a barrier, as well as repository software (60%, n=12). When the responses to questions about obstacles to quality were tabulated with the responses to quality rating, time limitations and staff hours came out as the top or joint-top answer, regardless of the quality rating. Finally, the authors present a sample of responses to the question on how metadata could be improved and these offer some solutions to staffing issues, the application of standards, and the repository system in use. Conclusion – The authors conclude that staffing, standards, and systems are all concerns in providing quality metadata. However, they suggest that standards and software issues could be overcome if adequate numbers of qualified staff are in place.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.101
metaresearch head score (Gemma)0.253
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Reporting · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.899
Threshold uncertainty score0.535

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1010.253
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0070.011
Science and technology studies0.0090.007
Scholarly communication0.0230.027
Open science0.0050.014
Research integrity0.0030.003
Insufficient payload (model declined to judge)0.0110.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.092
GPT teacher head0.377
Teacher spread0.285 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainReporting
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2016
Admission routes1
Has abstractyes

Explore more

Same venueEvidence Based Library and Information PracticeSame topicResearch Data Management PracticesFrench-language works237,207