MétaCan
Menu
Back to cohort
Record W1538444001 · doi:10.3414/me13-02-0029

Semi Automated Transformation to OWL Formatted Files as an Approach to Data Integration

2014· article· en· W1538444001 on OpenAlexaff
Adel Taweel, S. Miles, Yevgeniya Kovalchuk, Anastassia Spiridou, Benjamin Barratt, Uy Hoang, Siobhan Crichton, Brendan Delaney, C. Wolfe, Sheng‐Fu Liang

Bibliographic record

VenueMethods of Information in Medicine · 2014
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsInstitute of Population and Public Health
FundersNational Institutes of HealthNational Institute for Health and Care ResearchNIHR Biomedical Research Centre, Royal Marsden NHS Foundation Trust/Institute of Cancer Research
KeywordsComputer scienceMetadataInformation retrievalWeb Ontology LanguageOntologyInteroperabilityCorrectnessLinked dataMetadata repositoryWorld Wide WebSemantic WebDatabaseProgramming language

Abstract

fetched live from OpenAlex

INTRODUCTION: This article is part of the Focus Theme of METHODS of Information in Medicine on "Managing Interoperability and Complexity in Health Systems". BACKGROUND: Data heterogeneity is one of the critical problems in analysing, reusing, sharing or linking datasets. Metadata, whilst adding semantic description to data, adds an additional layer of complexity in the heterogeneity of metadata descriptors themselves. This can be managed by using a pre-defined model to extract the metadata, but this can reduce the richness of the data extracted. OBJECTIVES: to link the South London Stroke Register (SLSR), the London Air Pollution toolkit (LAP) and the Clinical Practice Research Datalink (CPRD) while transforming data into the Web Ontology Language (OWL) format. METHODS: We used a four-step transformation approach to prepare meta-descriptions, convert data, generate and update meta-classes and generate OWL files. We validated the correctness of the transformed OWL files by issuing queries and assessing results against the original source data. RESULTS: We have transformed SLSR LAP and CPRD into OWL format. The linked SLSR and CPRD OWL file contains 3644 male and 3551 female patients. The linked SLSR and LAP OWL file shows that there are 17 out of 35 outward postcode areas, where no overlapping data can support further analysis between SLSR and LAP. CONCLUSIONS: Our approach generated a resultant set of transformed OWL formatted files, which are in a query-able format to run individual queries, or can be easily converted into other more suitable formats for further analysis, and the transformation was faithful with no loss or anomalies. Our results have shown that the proposed method provides a promising general approach to address data heterogeneity.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.003
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.908
Threshold uncertainty score0.391

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.050
GPT teacher head0.414
Teacher spread0.364 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designOther design
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2014
Admission routes1
Has abstractyes

Explore more

Same venueMethods of Information in MedicineSame topicBiomedical Text Mining and OntologiesFrench-language works237,207