MétaCan
Menu
← Back to cohort
Record W4414189633 · doi:10.1101/2025.09.09.675119

Structure-based phylogenetic analysis reveals multiple events of convergent evolution of cysteine-rich antimicrobial peptides in legume-rhizobium symbiosis

2025· preprint· en· W4414189633 on OpenAlexafffund
Amira Boukherissa, Tatiana Timchenko, Mickaël Bourge, Peter Mergaert, George C. diCenzo, Jacqui A. Shykoff, Benoît Alunni, Ricardo C. Rodŕıguez de la Vega

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2025
Typepreprint
Languageen
FieldAgricultural and Biological Sciences
TopicLegume Nitrogen Fixing Symbiosis
Canadian institutionsQueen's University
FundersCentre National de la Recherche ScientifiqueMitacs
KeywordsSymbiosisPhylogenetic treePhylogeneticsCladeRhizobiaBacteriaConvergent evolutionSymbiotic bacteriaAntimicrobial peptides

Abstract

fetched live from OpenAlex

ABSTRACT Nitrogen is essential for plant growth, yet its availability often limits agricultural productivity. Some legumes have evolved a unique ability to form symbiotic relationships with nitrogen-fixing soil bacteria called rhizobia, enabling them to thrive in nitrogen-deficient soils. In five legume clades, an exploitive strategy has evolved in which rhizobia undergo Terminal Bacteroid Differentiation (TBD), where the bacteria become larger, polyploid, and have a permeabilized membrane. Terminally differentiated bacteria are associated with higher N 2 -fixation and, thus, a higher return on investment to the plant. In several members of the IRLC (Inverted Repeat-Lacking Clade) and the Dalbergioid clades of legumes, this differentiation process is triggered by a set of apparently unrelated plant antimicrobial peptides with membrane-damaging activity, known as Nodule-specific Cysteine-Rich (NCR) peptides. However, whether NCR peptides are also implicated in symbiotic TBD in other legume clades and whether they are evolutionarily related remains unknown. Here, to address the molecular identity of NCR peptides and their evolution in different legume clades, we performed inter- and intra-clade comparisons of NCR peptides in representative species of four TBD-inducing legume clades. First, we collected genomic and proteomic data of species for which NCR peptides are known (1523 NCR peptides). We then used sequence similarity-based clustering to regroup the NCR peptides, resulting in over 400 different NCR clusters, each clade-specific. We obtained Hidden Markov Models for each cluster and used them to predict NCR peptides in 21 legume genomes (6 clades), including newly generated deep-sequenced root and nodule RNA-seq data of Indigofera argentea (Indigoferoid clade) and newly assembled high-quality transcriptomes of Lupinus luteus and Lupinus mariae-josephae (Genistoid clade), using tailored gene prediction pipeline and transcriptome matching. This resulted in 3710 NCR peptides in species that induce TBD. To date, the rapid diversification of NCR peptides that reduces the sequence similarities has masked the origin of NCR peptide evolution. We obtained high-confidence structural models for one sequence of each cluster. We performed structure-based clustering and phylogenetics, which resulted in 23 superclusters (14 inter-clade and nine clade-specific) that we represent in a structural distance-based tree. Our study revealed that the evolution of NCR peptides is a mix of divergent and convergent processes within each clade. We further chose nine independently evolved NCR peptides to test in vitro whether they are functional analogs in symbiosis. Graphical abstract Overview of the experimental and computational workflow for NCR peptide detection, characterization, and structural analysis. Nodule and root samples from Indigofera argentea (8 weeks post-inoculation) were collected and subjected to RNA extraction, library preparation, and Illumina PE150 sequencing. Raw RNA-seq reads from two Lupinus species were also included ( Lupinus luteus and Lupinus mariae-josephae) . Bacteroid differentiation of I. argentea was assessed by flow cytometry and confocal microscopy. Transcriptomes were assembled de novo and analyzed for differential gene expression between root and nodule tissues. NCR peptides were identified from them and other legume genomes and transcriptomes using the SPADA pipeline and HMM profiles from NCR clusters of the known NCR peptides. The putative NCR peptides were filtered based on conserved cysteine motifs, length, and nodule expression to build an exhaustive NCR peptide database. 3D structural predictions of NCR clusters were performed using AlphaFold2 (pLDDT >70), followed by structural clustering (Foldseek) and phylogenetic analysis (Foldtree). Functional validation involved flow cytometry and antimicrobial assays (against Eschericha coli , Sinorhizobium meliloti , and Bacillus subtilis ), enabling structural and evolutionary characterization of NCR peptides. The green box at the top represents the experimental analysis, the blue box represents the sequence-based computational pipeline, the red box represents the structure-based computational pipeline, and the grey box at the bottom left represents the functional validation and interpretation of the results.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.002
Threshold uncertainty score0.003

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.000
Scholarly communication0.0010.000
Open science0.0000.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.009
GPT teacher head0.198
Teacher spread0.190 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)→Same topicLegume Nitrogen Fixing Symbiosis→French-language works237,207→