Marine biofilms: cyanobacteria factories for the global oceans
Bibliographic record
Abstract
Marine biofilms were newly revealed as a giant microbial diversity pool for global oceans. However, the cyanobacterial diversity in marine biofilms within the upper seawater column and its ecological and evolutionary implications remains undetermined. Here, we reconstructed a full picture of modern marine cyanobacteria habitats by re-analyzing 9.3 terabyte metagenomic data sets and 2,648 metagenome-assembled genomes (MAGs). The abundances of cyanobacteria lineages exclusively detected in marine biofilms were up to ninefold higher than those in seawater at similar sample size. Analyses revealed that cyanobacteria in marine biofilms are specialists with strong geographical and environmental constraints on their genome and functional adaption, which is in stark contrast to the generalistic features of seawater-derived cyanobacteria. Molecular dating suggests that the important diversifications in biofilm-forming cyanobacteria appear to coincide with the Great Oxidation Event (GOE), "boring billion" middle Proterozoic, and the Neoproterozoic Oxidation Event (NOE). These new insights suggest that marine biofilms are large and important cyanobacterial factories for the global oceans. IMPORTANCE: Cyanobacteria, highly diverse microbial organisms, play a crucial role in Earth's oxygenation and biogeochemical cycling. However, their connection to these processes remains unclear, partly due to incomplete surveys of oceanic niches. Our study uncovered significant cyanobacterial diversity in marine biofilms, showing distinct niche differentiation compared to seawater counterparts. These patterns reflect three key stages of marine cyanobacterial diversification, coinciding with major geological events in the Earth's history.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".