MétaCan
Menu
Retour à la cohorte
Enregistrement W2133243651 · doi:10.1093/biosci/biu126

Bioinformatics: Hypothesis Free—Or Hypotheses Freed?

2014· article· en· W2133243651 sur OpenAlexaff
Robert G. Beiko

Notice bibliographique

RevueBioScience · 2014
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueBioinformatics and Genomic Networks
Établissements canadiensDalhousie University
Organismes subventionnairesnon disponible
Mots-clésBiologyComputational biologyEvolutionary biology

Résumé

récupéré en direct d'OpenAlex

Bioinformatics articles contain ever-increasing numbers of statements about the ever-increasing volumes of data that are generated in the biosciences. Embedded in such “OMG, look at all these data!” statements are important points about the changing nature of biological research. Also embedded, however, is ambiguity about whether the work to follow is driven by the sheer availability of data or by an actual hypothesis to be tested. Further implied in these statements is the understanding that technological breakthroughs will sustain or increase the rate at which data accumulate. Whether this accumulation of data leads to greater insight is sometimes left as an exercise to the reader. In Life Out of Sequence: A Data-Driven History of Bioinformatics, author Hallam Stevens gives his readers a historical account of a field that began as an outgrowth of molecular biology—from a discipline of flasks and culture dishes, in which knowledge is built through careful, hypothesis-driven experimentation, to a technology-driven operation, in which data often precede hypotheses. Stevens is a historian of science whose background exemplifies some of the disciplinary shifts that have occurred in the history of bioinformatics. A key thread that Stevens documents in this book is the early role of recruits with a background in physics, such as Walter Goad and the author himself, who made the migratory transition to bioinformatics. Life Out of Sequence explores the history of bioinformatics in several complementary ways by including scientists (e.g., Goad, Margaret Dayhoff, James Ostell) who contributed to the emergence of the field, the ever-increasing role of data, the shaping and optimization of lab spaces and workflows, and the different ways in which visualization is used to translate massive amounts of data into a digestible and informative summary. Throughout Stevens's narrative, we read of milestones such as the prescient ideas and applications of the 1960s and 1970s, the struggle in the 1980s to bring both infrastructure and respect to bioinformatics, and the transformation of cell and molecular biology into “big science,” exemplified by the highly structured and optimized workflows at the Broad Institute. It is a compelling narrative that merges early successes with significant challenges to the field and its proponents, many of which persist to this day. Readers benefit from the book's extensive source material, as well as from dozens of interviews with scientists who have widely divergent views on bioinformatics. The opinions collected include Sydney Brenner's characterization of the field as “low input, high throughput, no output” (p. 66). Stevens also practices a form of embedded journalism by going “into the field” (i.e., the lab of Christopher Burge) to develop computer code in order to validate exon-shuffling events. The lab work described by the author provides a valuable example of the intersection of emerging technology, large-scale data analysis, and the statistics needed to distinguish real biological events from random noise. I found the discussion on ontologies to be particularly compelling. Anyone who has tried to describe gene lists in terms of putative functions or tried to cross-reference records using multiple databases will appreciate the limitations of bottom-up approaches that are driven by the views of individual researchers (sometimes without an overarching organizing principle). The limitations of such approaches quickly became evident as biology began to scale up and integrate data from multiple sources. The response has been the development of ontologies, top-down solutions that require consultation and negotiation among a wide variety of stakeholders. In documenting one of the earliest ontology projects in bioinformatics, the Gene Ontology consortium (GO; Ashburner et al. 2000), Stevens shows us the motivations, limitations, and opposition as Michael Ashburner and others tried to develop a hierarchical system with a controlled vocabulary to describe protein function. The success of GO is demonstrated through its primacy in many projects involving protein function (the 2000 paper introducing GO has been cited over 15,000 times), including the recent critical assessment of protein function annotation experiment, in which Radivojac and colleagues (2013) sought to predict GO terms for proteins with no known function, through the multitude of ontology projects that have emerged since the first release of GO. One challenge in writing a volume about bioinformatics is the definition of the term itself. Stevens first tackles this problem in the introduction and revisits the issue several times in the book. His interviews demonstrate pluralistic viewpoints: “Some understood [bioinformatics] as a limited set of tools for genome-centric biology, others as anything that had to do with biology and computers” (p. 43). Although it may be unrealistic to expect a single definition of bioinformatics to emerge from the book, the text quickly adopts a framework for “hypothesis-free,” “data-driven” science. The description of bioinformatics analyses (or experiments?) as being hypothesis free favors a view of the field in line with those expressed, for example, by Brenner. Whether or not the term is intended as a pejorative, it preempts the question of whether bread-and-butter techniques such as BLAST searching and genome-wide association studies are, in fact, hypothesis free: Are there not hypotheses embedded in the types of data that are collected and in the analytical methods that are chosen? Although no clear definition of bioinformatics is given, the book does contain an implicit definition that many biologists would reject as overly simplistic. The title of the concluding section, “The end of bioinformatics,” anticipates a world in which “its practices seem likely to become so ubiquitous that it will be absorbed into biology itself” (p. 219). But not every biologist will possess the computational and statistical wherewithal to push the discipline forward through the development of new algorithms and software to make sense of the data. Stevens portrays bioinformatics extensively as a new, data-driven discipline in contrast with molecular biology and its techniques that address one gene or system in considerable depth. Equating all of biology nearly exclusively to classic molecular biology inappropriately narrows the potential scope of the book, however. Although molecular biology revolutionized our understanding of the living world at the finest levels of organization, other disciplines are worth noting as influences to bioinformatics; the field would not exist as a useful discipline without the work of Karl Pearson or J. B. S. Haldane, for example. Indeed, the permutation tests used by Stevens to assign statistical significance to exon-shuffling events can claim a direct line of descent from R. A. Fisher. (There was also a missed opportunity to mention the intertwined development of ecology and statistics in the early 1900s, a revolution that parallels the later intertwining of biology and informatics documented in these pages.) Sequence-based bioinformatics is worthwhile in its own right, but the book presents a very restricted view of data and, more generally, of analytical techniques in biology. The historical account suffers as a result, because the analytical roots of bioinformatics go beyond sequence data and their representations. As a volume of history and ethnography, Life Out of Sequence documents the birth of a new discipline at the intersection of molecular biology and computer science. From this standpoint, it is a valuable contribution. The manner in which the loose definition of bioinformatics shapes the discussion, however, does a disservice to a field that extends beyond the processing of data. A more careful consideration of the scope of bioinformatics—and of its statistical roots—would have made for a much stronger and more balanced work.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,081
score de la tête « metaresearch » (Gemma)0,200
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Méthodes · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Commentaire · Signal consensuel: aucune
Score de désaccord entre enseignants0,919
Score d'incertitude au seuil0,431

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0810,200
Méta-épidémiologie (sens strict)0,0020,002
Méta-épidémiologie (sens large)0,0050,002
Bibliométrie0,0050,003
Études des sciences et des technologies0,0030,027
Communication savante0,0110,041
Science ouverte0,0070,007
Intégrité de la recherche0,0080,018
Charge utile insuffisante (le modèle a refusé de juger)0,0140,004

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,017
Tête enseignante GPT0,219
Écart entre enseignants0,201 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeThéorique ou conceptuel
DomaineMéthodes
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2014
Routes d'admission1
Résumé présentnon

Explorer davantage

Même revueBioScienceMême sujetBioinformatics and Genomic NetworksTravaux en français237 207