MétaCan
Menu
Back to cohort
Record W2899883143 · doi:10.3897/bdj.6.e29616

Incentivising use of structured language in biological descriptions: Author-driven phenotype data and ontology production

2018· article· en· W2899883143 on OpenAlexaff
Hong Cui, James Macklin, Joel L. Sachs, Anton A. Reznicek, Julian R. Starr, Bruce A. Ford, Lyubomir Penev, Hsin-Liang Chen

Bibliographic record

VenueBiodiversity Data Journal · 2018
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsUniversity of ManitobaUniversity of WinnipegUniversity of OttawaAgriculture and Agri-Food Canada
Fundersnot available
KeywordsVariation (astronomy)OntologyComputer scienceData scienceScale (ratio)Geography

Abstract

fetched live from OpenAlex

Phenotypes are used for a multitude of purposes such as defining species, reconstructing phylogenies, diagnosing diseases or improving crop and animal productivity, but most of this phenotypic data is published in free-text narratives that are not computable. This means that the complex relationship between the genome, the environment and phenotypes is largely inaccessible to analysis and important questions related to the evolution of organisms, their diseases or their response to climate change cannot be fully addressed. It takes great effort to manually convert free-text narratives to a computable format before they can be used in large-scale analyses. We argue that this manual curation approach is not a sustainable solution to produce computable phenotypic data for three reasons: 1) it does not scale to all of biodiversity; 2) it does not stop the publication of free-text phenotypes that will continue to need manual curation in the future and, most importantly, 3) It does not solve the problem of inter-curator variation (curators interpret/convert a phenotype differently from each other). Our empirical studies have shown that inter-curator variation is as high as 40% even within a single project. With this level of variation, it is difficult to imagine that data integrated from multiple curation projects can be of high quality. The key causes of this variation have been identified as semantic vagueness in original phenotype descriptions and difficulties in using standardised vocabularies (ontologies). We argue that the authors describing phenotypes are the key to the solution. Given the right tools and appropriate attribution, the authors should be in charge of developing a project's semantics and ontology. This will speed up ontology development and improve the semantic clarity of phenotype descriptions from the moment of publication. A proof of concept project on this idea was funded by NSF ABI in July 2017. We seek readers input or critique of the proposed approaches to help achieve community-based computable phenotype data production in the near future. Results from this project will be accessible through https://biosemantics.github.io/author-driven-production.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.086
metaresearch head score (Gemma)0.221
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Incentives · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.914
Threshold uncertainty score0.456

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0860.221
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0060.005
Science and technology studies0.0020.006
Scholarly communication0.0100.025
Open science0.0050.016
Research integrity0.0030.005
Insufficient payload (model declined to judge)0.0060.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.178
GPT teacher head0.332
Teacher spread0.154 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainIncentives
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2018
Admission routes1
Has abstractyes

Explore more

Same venueBiodiversity Data JournalSame topicBiomedical Text Mining and OntologiesFrench-language works237,207