High-coverage genomes of all extant penguin taxa
Bibliographic record
Abstract
Penguins (Sphenisciformes) are a highly diverse order of seabirds distributed widely across the Southern Hemisphere. They shared a common ancestor with Procellariiformes about 60 million years ago, and have subsequently transited from flying seabirds to flightless marine divers. Approximately 20 extant penguin species are recognized across six well-defined genera, ranging from the Galapagos Islands on the equator, to the oceanic temperate forests of New Zealand, the rocky coastlines of the sub-Antarctic islands, and the sea-ice around Antarctica. To inhabit such diverse and extreme environments, this speciose order has evolved many physiological and morphological adaptations. However, penguins are also highly sensitive to climate change, and most species are already declining or are predicted to decline under future climate change scenarios. Therefore, penguins are an exciting system for understanding the evolutionary processes of speciation, adaptation and demography. Genomic data are an emerging resource for addressing questions about such processes. Here we present 19 new high-coverage genomes that, together with two previously published genomes also available in GigaDB, encompass all extant penguin species. As such, this dataset provides a novel resource for understanding the evolutionary history within penguins, and between penguins and other avifauna. We believe that our dataset and project will be important for cultural heritage and the conservation of this iconic Southern Hemisphere species assemblage. Comparative and evolutionary genomic analyses are currently being carried out, and the consortium welcomes new members interested in contributing to this work. While this work is still underway we have published these 19 penguin genomes to provide early-release, while requesting researchers intending to use this data for similar cross-species comparisons to continue to follow the Fort Lauderdale and Toronto rules.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.064 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".