MétaCan
Menu
Back to cohort
Record W2133243651 · doi:10.1093/biosci/biu126

Bioinformatics: Hypothesis Free—Or Hypotheses Freed?

2014· article· en· W2133243651 on OpenAlexaff
Robert G. Beiko

Bibliographic record

VenueBioScience · 2014
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBioinformatics and Genomic Networks
Canadian institutionsDalhousie University
Fundersnot available
KeywordsBiologyComputational biologyEvolutionary biology

Abstract

fetched live from OpenAlex

Bioinformatics articles contain ever-increasing numbers of statements about the ever-increasing volumes of data that are generated in the biosciences. Embedded in such “OMG, look at all these data!” statements are important points about the changing nature of biological research. Also embedded, however, is ambiguity about whether the work to follow is driven by the sheer availability of data or by an actual hypothesis to be tested. Further implied in these statements is the understanding that technological breakthroughs will sustain or increase the rate at which data accumulate. Whether this accumulation of data leads to greater insight is sometimes left as an exercise to the reader. In Life Out of Sequence: A Data-Driven History of Bioinformatics, author Hallam Stevens gives his readers a historical account of a field that began as an outgrowth of molecular biology—from a discipline of flasks and culture dishes, in which knowledge is built through careful, hypothesis-driven experimentation, to a technology-driven operation, in which data often precede hypotheses. Stevens is a historian of science whose background exemplifies some of the disciplinary shifts that have occurred in the history of bioinformatics. A key thread that Stevens documents in this book is the early role of recruits with a background in physics, such as Walter Goad and the author himself, who made the migratory transition to bioinformatics. Life Out of Sequence explores the history of bioinformatics in several complementary ways by including scientists (e.g., Goad, Margaret Dayhoff, James Ostell) who contributed to the emergence of the field, the ever-increasing role of data, the shaping and optimization of lab spaces and workflows, and the different ways in which visualization is used to translate massive amounts of data into a digestible and informative summary. Throughout Stevens's narrative, we read of milestones such as the prescient ideas and applications of the 1960s and 1970s, the struggle in the 1980s to bring both infrastructure and respect to bioinformatics, and the transformation of cell and molecular biology into “big science,” exemplified by the highly structured and optimized workflows at the Broad Institute. It is a compelling narrative that merges early successes with significant challenges to the field and its proponents, many of which persist to this day. Readers benefit from the book's extensive source material, as well as from dozens of interviews with scientists who have widely divergent views on bioinformatics. The opinions collected include Sydney Brenner's characterization of the field as “low input, high throughput, no output” (p. 66). Stevens also practices a form of embedded journalism by going “into the field” (i.e., the lab of Christopher Burge) to develop computer code in order to validate exon-shuffling events. The lab work described by the author provides a valuable example of the intersection of emerging technology, large-scale data analysis, and the statistics needed to distinguish real biological events from random noise. I found the discussion on ontologies to be particularly compelling. Anyone who has tried to describe gene lists in terms of putative functions or tried to cross-reference records using multiple databases will appreciate the limitations of bottom-up approaches that are driven by the views of individual researchers (sometimes without an overarching organizing principle). The limitations of such approaches quickly became evident as biology began to scale up and integrate data from multiple sources. The response has been the development of ontologies, top-down solutions that require consultation and negotiation among a wide variety of stakeholders. In documenting one of the earliest ontology projects in bioinformatics, the Gene Ontology consortium (GO; Ashburner et al. 2000), Stevens shows us the motivations, limitations, and opposition as Michael Ashburner and others tried to develop a hierarchical system with a controlled vocabulary to describe protein function. The success of GO is demonstrated through its primacy in many projects involving protein function (the 2000 paper introducing GO has been cited over 15,000 times), including the recent critical assessment of protein function annotation experiment, in which Radivojac and colleagues (2013) sought to predict GO terms for proteins with no known function, through the multitude of ontology projects that have emerged since the first release of GO. One challenge in writing a volume about bioinformatics is the definition of the term itself. Stevens first tackles this problem in the introduction and revisits the issue several times in the book. His interviews demonstrate pluralistic viewpoints: “Some understood [bioinformatics] as a limited set of tools for genome-centric biology, others as anything that had to do with biology and computers” (p. 43). Although it may be unrealistic to expect a single definition of bioinformatics to emerge from the book, the text quickly adopts a framework for “hypothesis-free,” “data-driven” science. The description of bioinformatics analyses (or experiments?) as being hypothesis free favors a view of the field in line with those expressed, for example, by Brenner. Whether or not the term is intended as a pejorative, it preempts the question of whether bread-and-butter techniques such as BLAST searching and genome-wide association studies are, in fact, hypothesis free: Are there not hypotheses embedded in the types of data that are collected and in the analytical methods that are chosen? Although no clear definition of bioinformatics is given, the book does contain an implicit definition that many biologists would reject as overly simplistic. The title of the concluding section, “The end of bioinformatics,” anticipates a world in which “its practices seem likely to become so ubiquitous that it will be absorbed into biology itself” (p. 219). But not every biologist will possess the computational and statistical wherewithal to push the discipline forward through the development of new algorithms and software to make sense of the data. Stevens portrays bioinformatics extensively as a new, data-driven discipline in contrast with molecular biology and its techniques that address one gene or system in considerable depth. Equating all of biology nearly exclusively to classic molecular biology inappropriately narrows the potential scope of the book, however. Although molecular biology revolutionized our understanding of the living world at the finest levels of organization, other disciplines are worth noting as influences to bioinformatics; the field would not exist as a useful discipline without the work of Karl Pearson or J. B. S. Haldane, for example. Indeed, the permutation tests used by Stevens to assign statistical significance to exon-shuffling events can claim a direct line of descent from R. A. Fisher. (There was also a missed opportunity to mention the intertwined development of ecology and statistics in the early 1900s, a revolution that parallels the later intertwining of biology and informatics documented in these pages.) Sequence-based bioinformatics is worthwhile in its own right, but the book presents a very restricted view of data and, more generally, of analytical techniques in biology. The historical account suffers as a result, because the analytical roots of bioinformatics go beyond sequence data and their representations. As a volume of history and ethnography, Life Out of Sequence documents the birth of a new discipline at the intersection of molecular biology and computer science. From this standpoint, it is a valuable contribution. The manner in which the loose definition of bioinformatics shapes the discussion, however, does a disservice to a field that extends beyond the processing of data. A more careful consideration of the scope of bioinformatics—and of its statistical roots—would have made for a much stronger and more balanced work.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.081
metaresearch head score (Gemma)0.200
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.919
Threshold uncertainty score0.431

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0810.200
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0050.002
Bibliometrics0.0050.003
Science and technology studies0.0030.027
Scholarly communication0.0110.041
Open science0.0070.007
Research integrity0.0080.018
Insufficient payload (model declined to judge)0.0140.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.017
GPT teacher head0.219
Teacher spread0.201 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designTheoretical or conceptual
DomainMethods
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2014
Admission routes1
Has abstractno

Explore more

Same venueBioScienceSame topicBioinformatics and Genomic NetworksFrench-language works237,207