MétaCan
Menu
Back to cohort
Record W3151687099 · doi:10.29173/jchla29457

Challenges with organization, discoverability and access in Canadian open health data repositories

2021· article· en· W3151687099 on OpenAlexaffvenueabout
Gail M. Thornton, Ali Shiri

Bibliographic record

VenueJournal of the Canadian Health Libraries Association / Journal de l Association de bilbiothèques de la santé du Canada · 2021
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsDiscoverabilityMetadataComputer scienceWorld Wide WebMetadata repositoryOpen dataInteroperabilityInformation retrievalData elementData science

Abstract

fetched live from OpenAlex

Introduction: Open health data provides healthcare professionals, biomedical researchers and the general public with access to health data which has the potential to improve healthcare delivery and policy. The challenge is to create and implement appropriate metadata, or structured data about the data, to ensure that data are easy to discover, access and re-use. The goal of this study is to identify, evaluate and compare Canadian open health data repositories for their searching, browsing and navigation functionalities, the richness of their metadata description practices, and their metadata-based filtering mechanisms. Methods: Metadata-based search and browsing was evaluated in addition to the number and nature of metadata elements. Six Canadian open health data repositories across national, provincial and institutional levels were evaluated. Data collected using verbatim text recording was evaluated using an analytical framework based on the 2019 Dataverse North Metadata Best Practices guide and 2019 Data Citation Implementation Project roadmap. Results: All repositories required filtering to access "open health data." All repositories included 'subject' facets for filtering, and 'title' and 'description' on the Results List. Use case evaluations suggest improvements including advanced search, health-specific search terms, records for all repositories, and links to related publications. Discussion: Consistent use of 'title' and 'description' suggests that an interoperable interface is possible. Inconsistencies in records indicate the need for explicit, easy to find mechanisms to access metadata in repositories. The analytical framework represents first draft guidelines for metadata creation and implementation to improve organization, discoverability, and access to Canadian open health data.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.153
metaresearch head score (Gemma)0.305
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Scholarly communication, Open science
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.989
Threshold uncertainty score0.939

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1530.305
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0330.058
Science and technology studies0.0310.021
Scholarly communication0.0510.022
Open science0.0110.022
Research integrity0.0040.003
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.037
GPT teacher head0.338
Teacher spread0.301 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designQualitative
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2021
Admission routes3
Has abstractyes

Explore more

Same venueJournal of the Canadian Health Libraries Association / Journal de l Association de bilbiothèques de la santé du CanadaSame topicResearch Data Management PracticesFrench-language works237,207