Getting clear about the F-word in genomics
Notice bibliographique
Résumé
Although biology is generally awash with adaptationist "just-so" stories, the situation in molecular biology and genomics is particularly bad.Various types of non-coding DNA are routinely interpreted as functional without adequate consideration of non-adaptationist alternative hypotheses [1].Part of the problem is surely due to a failure in these disciplines to appreciate theoretical developments in population genetics, which outline the conditions under which genetic elements are selected [2].However, as a number of authors have noted, the problem is also partly due to a confusion about the various possible meanings of "function" in biology [3][4][5].Our central thesis is that there exists an overlooked dichotomy in the way that researchers see natural selection to be related to function.Traits or genetic elements that are merely under purifying selection have what we call maintenance functions whereas those that have historically been under directional selection have origin functions.We argue that ignoring this distinction encourages a form of pan-adaptationism, where highly plausible non-adaptive explanations for the origins of certain genetic elements or traits are themselves ignored.Thus, our recommendation is for researchers to always clarify which sense of "function" they mean (origin or maintenance) when talking or writing about selected effects.Before developing this argument, it is important to clarify our position by distinguishing selection-based notions of functions as a class from causal role (CR) functions.Although this distinction is widely recognized by philosophers of biology our sense is that it remains unfamiliar to many biologists.The CR definition of function is extremely permissive.It applies to any of the effects which a component has on the system(s) that contain it, irrespective of their impact (or that system's impact) on fitness.For example, a mobile genetic element which elevates mutation rate in the genome has this effect as one of its CR functions, even if it causes a net decrease in organismal fitness.Such permissiveness in the definition of CR function has led some researchers to dismiss this concept [6].This reaction is understandable when it comes from researchers working in the disciplines of ecology or evolution, where there is often an emphasis on the ecological roles performed by a given trait and their effects on organismal fitness.More controversial is whether researchers working in molecular biology or bioinformatics would embrace the CR concept once its commitment to fitness neutrality is made explicit.On the one hand, investigators in these disciplines might point out that they use methods (e.g.biochemical interaction measurements) that can only establish an entity's causal roles.To infer a contribution to fitness (and thus selection) requires an additional and difficult-to-prove inference, namely that those causal role "functions" have indeed been under selection.As it turns out, sometimes those inferences are poorly supported-as in the publicity surrounding ENCODE, which we discuss below.Nonetheless, from this perspective it makes sense to view much of the work in molecular biology or bioinformatics as being focussed primarily on CR functions.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,014 | 0,042 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,002 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,005 | 0,039 |
| Communication savante | 0,008 | 0,028 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,010 | 0,023 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,012 | 0,007 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».