The Amazon River microbiome, a story of humic carbon
Bibliographic record
Abstract
Abstract The Amazon River basin sustains dramatic hydrochemical gradients defined by three water types: white, clear and black waters. Black waters contain important loads of allochthonous humic dissolved organic carbon (DOC), mostly coming from bacteria-mediated lignin degradation, a process that remains understudied. Here, we identified the main bacterial taxa and functions associated with contrasting Amazonian water types, and shed light on their potential implication in the lignin degradation process. We performed an extensive field bacterioplankton sampling campaign from the three Amazonian water types, and combined our observations to a meta-analysis of 90 Amazonian basin shotgun metagenomes used to build a tailored functional inference database. We showed that the overall quality of DOC is a major driver of bacterioplankton structure, transcriptional activity and functional repertory. We also showed that among the taxa mostly associated to differences between water types, Polynucleobacter sinensis particularly stood out, as its abundance and transcriptional activity was strongly correlated to black water environments, and specially to humic DOC concentration. Screening the reference genome of this bacteria, we found genes coding for enzymes implicated in all the main lignin degradation steps, suggesting that this bacteria may play key roles in the carbon cycle processes within the Amazon basin.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".