Getting clear about the F-word in genomics
Bibliographic record
Abstract
Although biology is generally awash with adaptationist "just-so" stories, the situation in molecular biology and genomics is particularly bad.Various types of non-coding DNA are routinely interpreted as functional without adequate consideration of non-adaptationist alternative hypotheses [1].Part of the problem is surely due to a failure in these disciplines to appreciate theoretical developments in population genetics, which outline the conditions under which genetic elements are selected [2].However, as a number of authors have noted, the problem is also partly due to a confusion about the various possible meanings of "function" in biology [3][4][5].Our central thesis is that there exists an overlooked dichotomy in the way that researchers see natural selection to be related to function.Traits or genetic elements that are merely under purifying selection have what we call maintenance functions whereas those that have historically been under directional selection have origin functions.We argue that ignoring this distinction encourages a form of pan-adaptationism, where highly plausible non-adaptive explanations for the origins of certain genetic elements or traits are themselves ignored.Thus, our recommendation is for researchers to always clarify which sense of "function" they mean (origin or maintenance) when talking or writing about selected effects.Before developing this argument, it is important to clarify our position by distinguishing selection-based notions of functions as a class from causal role (CR) functions.Although this distinction is widely recognized by philosophers of biology our sense is that it remains unfamiliar to many biologists.The CR definition of function is extremely permissive.It applies to any of the effects which a component has on the system(s) that contain it, irrespective of their impact (or that system's impact) on fitness.For example, a mobile genetic element which elevates mutation rate in the genome has this effect as one of its CR functions, even if it causes a net decrease in organismal fitness.Such permissiveness in the definition of CR function has led some researchers to dismiss this concept [6].This reaction is understandable when it comes from researchers working in the disciplines of ecology or evolution, where there is often an emphasis on the ecological roles performed by a given trait and their effects on organismal fitness.More controversial is whether researchers working in molecular biology or bioinformatics would embrace the CR concept once its commitment to fitness neutrality is made explicit.On the one hand, investigators in these disciplines might point out that they use methods (e.g.biochemical interaction measurements) that can only establish an entity's causal roles.To infer a contribution to fitness (and thus selection) requires an additional and difficult-to-prove inference, namely that those causal role "functions" have indeed been under selection.As it turns out, sometimes those inferences are poorly supported-as in the publicity surrounding ENCODE, which we discuss below.Nonetheless, from this perspective it makes sense to view much of the work in molecular biology or bioinformatics as being focussed primarily on CR functions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.042 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.005 | 0.039 |
| Scholarly communication | 0.008 | 0.028 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.010 | 0.023 |
| Insufficient payload (model declined to judge) | 0.012 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".