Bibliographic record
Abstract
Abstract The first complete bacterial genome sequence, for Haemophilus influenzae, was published in 1995,1 followed the same year by the sequence for Mycoplasma genitalium. This simple, wall-less organism has one of the smallest genomes, only 580,070 base pairs and carrying just 470 genes. The genome sequence of a cyanobacterium (formerly called a blue-green alga) came in 1996, and in 1997 those for three chief model microbes: Escherichia coli, the workhorse of molecular microbiology; Bacillus subtilis, the model gram-positive bacillus; and baker’s yeast, Saccharomyces cerevisiae, a favorite for studying eukaryote molecular and cell biology. Completion of the yeast sequence was a triumph of organization because it resulted from a cooperation among100 laboratories in Canada, Japan, the United States, and Europe, including a European Union consortium of laboratories from almost every member state. The article in Nature summarizing the results had what may be a record of 633 authors.2 It was by far the largest genome to be completed, with more than 12 mil- lion base pairs of DNA distributed among 16 separate chromosomes, although encoding only about 6000 genes. This is far fewer than in a bacterial chromosome of the same size, because eukaryotic chromosomes contain long stretches of DNA that do not code for proteins. By early 2006, there were 316 complete genomes of bacteria alone, and another 933 were in progress.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.098 | 0.070 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".