MétaCan
Menu
Back to cohort
Record W4410350959 · doi:10.1101/2025.05.07.652750

Improving plant functional annotation from knowledge graphs using Graph Neural Networks

2025· preprint· en· W4410350959 on OpenAlexaff
Tran Gia Bao Ngo, Christophe Liseron-Monfils, Jordan Ubbens, Paula Ashe, David Konkin

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2025
Typepreprint
Languageen
FieldComputer Science
TopicData Mining Algorithms and Applications
Canadian institutionsNational Research Council Canada
Fundersnot available
KeywordsKnowledge graphComputer scienceAnnotationArtificial neural networkGraphArtificial intelligenceTheoretical computer science

Abstract

fetched live from OpenAlex

Abstract Annotating genes is essential to crop development, and understanding gene functions sheds light on developing crop improvement strategies, such as marker-assisted breeding, genetic modification, or pest resistance. Through an extensive experimental effort and computational annotation projection, tens of thousands of genes have been annotated across plant species, with most of the gene annotations focusing on a well-studied species, Arabidopsis thaliana, but this represents a small fraction of the hundreds of thousands of genes across these different plant species. Phenotypes and their traits result from multiple processes and events involving multiscale information encoded from different omics, such as genomes, proteomes, or transcriptomes. This stresses a need for an efficient computational approach to capture and integrate information from biological networks and transfer this knowledge from well-studied species to unknown species to annotate and discover functional relationships between annotations and genes. Despite recent progress, existing methods only consider one or a few omics levels to perform reasoning on functional annotation-to-gene relations. The main objective of this study is to generate and explore a large-scale plant biological knowledge graph, the DasDB, and to enrich gene functional annotation linked to genes in different species using graph neural networks (GNNs). Integrating various data sources from different omics has resulted in a comprehensive graph database, facilitating researchers’ in-depth understanding of complex biological networks at the highest level. In addition, applying GNNs on a large-scale knowledge graph database has shown promise in the ability of deep learning models to transfer this information from well-studied plant species to less-characterized plant species, outperforming the transfer of information done using only orthology relationships. This study benchmarks a new research direction in producing new functional annotation discovery in plant species with limited functional annotations. This pipeline was applied to a specific research problem: the mechanism involved in pea nodule nitrogen fixation. We identified known gene markers of this process through a systematic analysis of the DasDB, showing the relevance of our approach. Furthermore, new potential targets to better understand and improve this process were identified.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.810
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.228
Teacher spread0.207 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicData Mining Algorithms and ApplicationsFrench-language works237,207