Bibliographic record
Abstract
The publication of the E. coli O157:H7 genome sequence last year should have been the last straw on the back of that camel we call ‘the bacterial species’. The hamburger pathogen carries 1387 genes absent from the genome of its more-or-less benign ‘conspecific’, K12, which itself bears 528 genes missing from O157:H7. The differences are by no means confined to a few large pathogenicity islands: Perna and collaborators (Perna et al., 2001) counted 177 ‘O islands’ unique to O157:H7, and reciprocally 234 ‘K islands’, unique to K12. A similar picture appears when we look at the recently published genomes of two strains of Salmonella enterica. One of these is that same ‘S. typhimurium’ (S. enterica serovar Typhimurium LT2) whose genetic map had long been known to be remarkably similar to that of Escherichia coli. A whopping 29% of LT2’s genes are missing from K12, and 11% are missing from the conspecific S. enterica serovar Typhi CT18 (McClelland et al., 2001). It is possible to reconcile similar maps with dissimilar gene contents, using Roger Milkman’s clonal frame model, developed well before there were whole genome sequences to compare. This embraces the fact that bacterial genomes change mostly through the insertion or deletion (by legitimate or illegitimate recombination) of DNA segments of a few hundred basepairs to a few dozen genes, without disrupting overall genome order. The surprise is not that we find strain-specific genes or islands, but that we find so amazingly many, and that so many seem to have been acquired illegitimately (by lateral gene transfer)! Indeed, it is now clearly inappropriate to consider any single sequenced prokaryotic genome as the genome of its species. Last year, Lan and Reeves (Lan and Reeves, 2000) elaborated a more realistic concept of a ‘species genome’, a virtual entity encompassing all genes found in known strains of a single species, and comprising two fractions: a ‘core’ of genes present in > 95% of strains and an ‘auxiliary set’ found in 1–95%. Surely we must accept something like this, while at the same time admitting that the traits that help us decide which strains belong to which species are not essentially different in terms of their fugacity from those we relegate to the auxiliary set. One might suppose that such instability in gene content is particularly common in human (and veterinary) pathogens, because of the growth and mobility of our own population and because of our attempts to defend ourselves with antibiotics. But unperturbed and perturbed ‘natural’ environments can surely impose changes in microbial-selective regimens that are just as sudden and profound as any experienced in our cities or hospitals. Furthermore, the many specialized genetic structures that promote within- and between-species recombination and gene transfer (phages, plasmids, mating and transformation systems, transposons, integrons, mobile gene cassettes and the like) all pre-date our own appearance as hosts and opponents in a pharmaceutical arms race. (In fact, to the extent that such genetic agents have been operating across species boundaries for millions of millennia, they populate a fourth evolutionary ‘domain of Life’, albeit not an organismal one). It seems prudent to imagine that once we have many sequenced strains of Thermotoga maritima or Sulfolobus solfataricus or even Nostoc punctiforme, we will be equally hard pressed to say which represents the genome of its species. Genome sequencing, especially the highly informative but not quite complete sequencing practised by the Joint Genome Initiative (www.jgi.gov/JGI_microbial/html/) gets ever faster and cheaper, but still no one is likely to sequence hundreds of Thermotogas soon. More economical methods will be needed if we are to get a grip on within-species genomic diversity. Suppressive subtractive hybridization and hybridization to microarrays of genes from sequenced representatives have already seen used for this purpose, and sequencing of multiple environmental BAC clones that bear identical (or nearly identical) SSU rRNA genes offers another window. Our own experience with the first method suggests that it provides an efficient way of finding differences between genomes that are very close (in terms of the sequences of the genes they do share). But as this parameter decreases there will be more ‘false positives’, and the discovery of true positives seems less exciting: we are after all not surprised if strains of different species have different genes. Indeed the effort to define species genomes may force us to confront, for once and all, the fact that there may never be a coherent species concept for prokaryotes. Such an intellectual construct is, of course, invaluable for identification and classification. Purely operational definitions based on SSU rRNA sequence similarity or genomic hybridization, or a combination of these and multilocus sequence typing, have essential roles to play in organizing the science of microbiology. But this doesn’t mean that Nature has organized herself along species lines. The most appealing species concept in non-microbial biology, the ‘biological species concept’ manifestly does not apply to prokaryotes. Either they have sex too infrequently (and are thus like problematic ‘asexual species’ of animals) or they have it with partners that are too remotely related (and thus all comprise together one eternal global species). There are to be sure species-like behaviours (shared gene pools with boundaries, however leaky) among some prokaryotes, but no consistent mapping of such behaviours to taxonomic units defined phenotypically. In fact, we have a good general understanding of bacterial population genetics and genome evolution now: it’s just the attempt to reconcile this knowledge with our idea that there should be ‘species’ that causes trouble. As Marcel Duchamp said, in some totally different context: ‘There is no solution because there is no problem’. Although claims that we can cultivate fewer than 1% of prokaryotes may founder on the indefinability (and thus uncountability) of species, there is no question that there is enormous uncharacterized diversity out there, as assessed by phylotyping with SSU rRNA and other molecular markers. In fact, the ‘species genome’ phenomenon tell us that the true enormity of microbial diversity has yet to be appreciated. How many different genomes (differing in gene content) are hiding behind a single environmental SSU rRNA phylotype, and how many under the phylotype clusters are commonly found? Is the iceberg of diversity, of which the cultivatable 1% is just the tip, itself just the tip of a still larger iceberg, extending orthogonally from phylotype space into ‘genomo-type’ space? When we can map the smaller diversity into the larger, will there still be regions of denser occupation that correspond to currently recognized taxa, from domains to strains? Or will the overlaps by so extensive that we abandon this way of looking at the microbial world altogether?
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.011 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".