MétaCan
Menu
Back to cohort
Record W3014902736 · doi:10.1371/journal.pgen.1008702

Getting clear about the F-word in genomics

2020· article· en· W3014902736 on OpenAlexafffund
Stefan Linquist, W. Ford Doolittle, Alexander F. Palazzo

Bibliographic record

VenuePLoS Genetics · 2020
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicRNA and protein synthesis mechanisms
Canadian institutionsUniversity of TorontoDalhousie UniversityUniversity of Guelph
FundersSocial Sciences and Humanities Research Council of CanadaNatural Sciences and Engineering Research Council of Canada
KeywordsBiologyGenomicsWord (group theory)Computational biologyGeneticsEvolutionary biologyGenomeLinguisticsGene

Abstract

fetched live from OpenAlex

Although biology is generally awash with adaptationist "just-so" stories, the situation in molecular biology and genomics is particularly bad.Various types of non-coding DNA are routinely interpreted as functional without adequate consideration of non-adaptationist alternative hypotheses [1].Part of the problem is surely due to a failure in these disciplines to appreciate theoretical developments in population genetics, which outline the conditions under which genetic elements are selected [2].However, as a number of authors have noted, the problem is also partly due to a confusion about the various possible meanings of "function" in biology [3][4][5].Our central thesis is that there exists an overlooked dichotomy in the way that researchers see natural selection to be related to function.Traits or genetic elements that are merely under purifying selection have what we call maintenance functions whereas those that have historically been under directional selection have origin functions.We argue that ignoring this distinction encourages a form of pan-adaptationism, where highly plausible non-adaptive explanations for the origins of certain genetic elements or traits are themselves ignored.Thus, our recommendation is for researchers to always clarify which sense of "function" they mean (origin or maintenance) when talking or writing about selected effects.Before developing this argument, it is important to clarify our position by distinguishing selection-based notions of functions as a class from causal role (CR) functions.Although this distinction is widely recognized by philosophers of biology our sense is that it remains unfamiliar to many biologists.The CR definition of function is extremely permissive.It applies to any of the effects which a component has on the system(s) that contain it, irrespective of their impact (or that system's impact) on fitness.For example, a mobile genetic element which elevates mutation rate in the genome has this effect as one of its CR functions, even if it causes a net decrease in organismal fitness.Such permissiveness in the definition of CR function has led some researchers to dismiss this concept [6].This reaction is understandable when it comes from researchers working in the disciplines of ecology or evolution, where there is often an emphasis on the ecological roles performed by a given trait and their effects on organismal fitness.More controversial is whether researchers working in molecular biology or bioinformatics would embrace the CR concept once its commitment to fitness neutrality is made explicit.On the one hand, investigators in these disciplines might point out that they use methods (e.g.biochemical interaction measurements) that can only establish an entity's causal roles.To infer a contribution to fitness (and thus selection) requires an additional and difficult-to-prove inference, namely that those causal role "functions" have indeed been under selection.As it turns out, sometimes those inferences are poorly supported-as in the publicity surrounding ENCODE, which we discuss below.Nonetheless, from this perspective it makes sense to view much of the work in molecular biology or bioinformatics as being focussed primarily on CR functions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.042
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.014
Threshold uncertainty score0.074

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0140.042
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0020.002
Science and technology studies0.0050.039
Scholarly communication0.0080.028
Open science0.0020.003
Research integrity0.0100.023
Insufficient payload (model declined to judge)0.0120.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.025
GPT teacher head0.230
Teacher spread0.205 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations33
Published2020
Admission routes2
Has abstractyes

Explore more

Same venuePLoS GeneticsSame topicRNA and protein synthesis mechanismsFrench-language works237,207