Comparative Phylogenomics and Phylotranscriptomics Provide Insights into the Genetic Complexity of Nitrogen Fixing Root Nodule Symbiosis
Bibliographic record
Abstract
Abstract Plant root nodule symbiosis (RNS) with mutualistic nitrogen-fixing bacteria is restricted to a single clade of angiosperms, the Nitrogen-Fixing Nodulation Clade (NFNC), and is best understood in the legume family. It is widely accepted that nodulation originated through the assembly of modules recruited from existing functions, such as mycorrhizal symbiosis, polar growth, and lateral root development. Because nodulating species are scattered within the NFNC, the number of times nodulation has evolved or has been lost has been a matter of considerable speculation. This interesting evolutionary question has practical implications concerning the ease with which nodulation might be engineered in non-nodulating crop plants. Nodulating species share many commonalities, due either to divergence from a common ancestor over 100 million years ago or to convergence or deep homology following independent origins over that same time period. In either case, comparative analyses of diverse nodulation syndromes can provide insights into constraints on nodulation—what must be acquired or cannot be lost for a functional symbiosis—and what the latitude is for variation in the symbiosis. However, much remains to be learned about nodulation, especially outside of legumes. Here we present new information across the spectrum of nodulating groups. We find no evidence for convergence at the level of amino acid residues or gene family expansion across the NFNC. Our phylogenomic analyses further emphasize the uniqueness of the transcription factor, NIN, as a master regulator of nodulation, and identify key mutations affecting its function across the NFNC. We find that nodulation genes are over-represented among orthologous gene groups (OGs) present in the NFNC common ancestor, but that lineage-specific OGs play major roles in nodulation. We identified over 900,000 conserved noncoding elements (CNEs), of which over 300,000 were unique to NFNC species. A significant proportion of these are associated with nodulation-related genes and thus are candidates for transcriptional regulators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".