CLCNet: a contrastive learning and chromosome-aware network for genomic prediction in plants
Notice bibliographique
Résumé
Abstract Genomic selection (GS) leverages genome-wide markers and phenotypes to predict breeding values, with its effectiveness largely dependent on the accuracy of genomic prediction (GP) models. However, GP methods often struggle to capture inter-individual variability and are limited by the curse of dimensionality, where the number of SNPs far exceeds the sample size. To address these challenges, we present CLCNet (Contrastive Learning and Chromosome-aware Network), a novel deep learning framework that integrates contrastive learning and chromosome-aware feature modeling. CLCNet comprises two key components: (i) a contrastive learning module that enhances the model’s ability to capture fine-grained, genotype-dependent phenotypic differences among individuals, and (ii) a chromosome-aware module that captures structured feature selection at both chromosome and genome levels, thereby distilling the most informative SNPs. We evaluated CLCNet across four crop species, covering ten agronomically important traits, and compared it with a diverse set of classical linear, machine learning, and deep learning models. CLCNet achieved superior prediction performance, with statistically significant improvements in Pearson correlation coefficient (PCC), ranging from 0.34% to 12.19% over baseline, together with reduced mean squared error (MSE). Performance gains were more pronounced for traits with moderate linkage disequilibrium (LD; r 2 = 0.21-0.36) and high heritability ( h 2 > 0.66), such as those in maize, rapeseed, and soybean. For cotton traits characterized by high LD ( r 2 = 0.74) and lower heritability ( h 2 < 0.50), CLCNet maintained robust performance without degradation. Overall, these results demonstrate that CLCNet is an effective framework for improving genomic prediction accuracy and holds strong potential for practical applications in plant breeding. Short abstract CLCNet is a novel deep learning framework for genomic prediction that integrates contrastive learning with chromosome-aware feature selection. By jointly modeling inter-individual genotype–phenotype variation and chromosomal genomic structure, CLCNet improves prediction accuracy under high-dimensional, low-sample-size conditions. Across four crop species and ten agronomic traits, CLCNet consistently outperformed classical statistical, machine learning, and existing deep learning models. The framework also identified biologically relevant SNPs and candidate genes, demonstrating its potential for practical applications in genomic selection and computational plant breeding. Key points We propose CLCNet, a multi-task deep learning framework that integrates contrastive learning with chromosome-aware feature selection for genomic prediction, under high-dimensional, low-sample-size conditions. The chromosome-aware module explicitly exploits genomic structural information to select representative and informative SNPs across chromosomes. Contrastive learning improves model robustness by stabilizing representation learning and reducing the influence of random effects across samples. By complementing GWAS analyses, CLCNet provides additional insights into genotype–phenotype relationships with potential relevance for gene discovery. Biographical Note Jiangwei Huang is a PhD candidate at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. His research focuses on genomic prediction, deep learning, and computational plant breeding. Zhihan Yang is a PhD candidate at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. Her research interests include genomic prediction and bioinformatics. Rongcheng Han is an associate professor at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. His research focuses on bioinformatics and plant phenomics. Yuqiang Jiang is a professor at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. His research interests include plant genomics, plant phenomics and genetic improvement. Organization description The Institute of Genetics and Developmental Biology, Chinese Academy of Sciences, is a leading research institute focusing on genetics, genomics, molecular breeding, bioinformatics, and systems biology in plants and animals.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».