Improving plant functional annotation from knowledge graphs using Graph Neural Networks
Notice bibliographique
Résumé
Abstract Annotating genes is essential to crop development, and understanding gene functions sheds light on developing crop improvement strategies, such as marker-assisted breeding, genetic modification, or pest resistance. Through an extensive experimental effort and computational annotation projection, tens of thousands of genes have been annotated across plant species, with most of the gene annotations focusing on a well-studied species, Arabidopsis thaliana, but this represents a small fraction of the hundreds of thousands of genes across these different plant species. Phenotypes and their traits result from multiple processes and events involving multiscale information encoded from different omics, such as genomes, proteomes, or transcriptomes. This stresses a need for an efficient computational approach to capture and integrate information from biological networks and transfer this knowledge from well-studied species to unknown species to annotate and discover functional relationships between annotations and genes. Despite recent progress, existing methods only consider one or a few omics levels to perform reasoning on functional annotation-to-gene relations. The main objective of this study is to generate and explore a large-scale plant biological knowledge graph, the DasDB, and to enrich gene functional annotation linked to genes in different species using graph neural networks (GNNs). Integrating various data sources from different omics has resulted in a comprehensive graph database, facilitating researchers’ in-depth understanding of complex biological networks at the highest level. In addition, applying GNNs on a large-scale knowledge graph database has shown promise in the ability of deep learning models to transfer this information from well-studied plant species to less-characterized plant species, outperforming the transfer of information done using only orthology relationships. This study benchmarks a new research direction in producing new functional annotation discovery in plant species with limited functional annotations. This pipeline was applied to a specific research problem: the mechanism involved in pea nodule nitrogen fixation. We identified known gene markers of this process through a systematic analysis of the DasDB, showing the relevance of our approach. Furthermore, new potential targets to better understand and improve this process were identified.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».