MétaCan
Menu
Back to cohort
Record W2207671015 · doi:10.1055/s-0038-1638748

Key Concepts to Assess the Readiness of Data for International Research: Data Quality, Lineage and Provenance, Extraction and Processing Errors, Traceability, and Curation

2011· article· en· W2207671015 on OpenAlexaff
Simon de Lusignan, Siaw‐Teng Liaw, Paul Krause, Vasa Ćurčin, Georgios Michalakidis, L. Agreus, P. Leysen, Nicola Shaw, Kumara Mendis

Bibliographic record

VenueYearbook of Medical Informatics · 2011
Typearticle
Languageen
FieldDecision Sciences
TopicScientific Computing and Data Management
Canadian institutionsAlgoma UniversityEsri (Canada)
FundersDirectorate-General for Information Society and MediaEuropean Commission
KeywordsMetadataComputer scienceData qualityData scienceData curationMetadata repositoryData warehouseTraceabilityData extractionInformation retrievalData elementDatabaseWorld Wide WebEngineering

Abstract

fetched live from OpenAlex

OBJECTIVE: To define the key concepts which inform whether a system for collecting, aggregating and processing routine clinical data for research is fit for purpose. METHODS: Literature review and shared experiential learning from research using routinely collected data. We excluded socio-cultural issues, and privacy and security issues as our focus was to explore linking clinical data. RESULTS: Six key concepts describe data: (1) DATA QUALITY: the core Overarching concept - Are these data fit for purpose? (2) Data provenance: defined as how data came to be; incorporating the concepts of lineage and pedigree. Mapping this process requires metadata. New variables derived during data analysis have their own provenance. (3) Data extraction errors and (4) Data processing errors, which are the responsibility of the investigator extracting the data but need quantifying. (5) Traceability: the capability to identify the origins of any data cell within the final analysis table essential for good governance, and almost impossible without a formal system of metadata; and (6) Curation: storing data and look-up tables in a way that allows future researchers to carry out further research or review earlier findings. CONCLUSION: There are common distinct steps in processing data; the quality of any metadata may be predictive of the quality of the process. Outputs based on routine data should include a review of the process from data origin to curation and publish information about their data provenance and processing method.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.482
metaresearch head score (Gemma)0.634
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.518
Threshold uncertainty score0.639

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4820.634
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0310.032
Science and technology studies0.0070.035
Scholarly communication0.0310.045
Open science0.0060.018
Research integrity0.0070.007
Insufficient payload (model declined to judge)0.0060.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.814
GPT teacher head0.618
Teacher spread0.197 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations48
Published2011
Admission routes1
Has abstractyes

Explore more

Same venueYearbook of Medical InformaticsSame topicScientific Computing and Data ManagementFrench-language works237,207