MétaCan
Menu
Back to cohort
Record W4405975676 · doi:10.1093/bib/bbag270

CLCNet: a contrastive learning and chromosome-aware network for genomic prediction in plants

2024· preprint· en· W4405975676 on OpenAlexfundno aff
Zhihan Yang, Mou Yin, Chao Li, Jinmin Li, Yu Wang, Lu Huang, Miaomiao Li, Chengzhi Liang, Fei He, Runhao Han, Yuqiang Jiang

Bibliographic record

VenueBriefings in Bioinformatics · 2024
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic Mapping and Diversity in Plants and Animals
Canadian institutionsnot available
FundersInstitute of GeneticsNational Key Research and Development Program of ChinaPeking University Health Science CenterPeking UniversityChinese Academy of Sciences
KeywordsArtificial intelligenceEpistasisMachine learningFeature selectionComputer scienceBiologyGeneticsGene

Abstract

fetched live from OpenAlex

Abstract Genomic selection (GS) leverages genome-wide markers and phenotypes to predict breeding values, with its effectiveness largely dependent on the accuracy of genomic prediction (GP) models. However, GP methods often struggle to capture inter-individual variability and are limited by the curse of dimensionality, where the number of SNPs far exceeds the sample size. To address these challenges, we present CLCNet (Contrastive Learning and Chromosome-aware Network), a novel deep learning framework that integrates contrastive learning and chromosome-aware feature modeling. CLCNet comprises two key components: (i) a contrastive learning module that enhances the model’s ability to capture fine-grained, genotype-dependent phenotypic differences among individuals, and (ii) a chromosome-aware module that captures structured feature selection at both chromosome and genome levels, thereby distilling the most informative SNPs. We evaluated CLCNet across four crop species, covering ten agronomically important traits, and compared it with a diverse set of classical linear, machine learning, and deep learning models. CLCNet achieved superior prediction performance, with statistically significant improvements in Pearson correlation coefficient (PCC), ranging from 0.34% to 12.19% over baseline, together with reduced mean squared error (MSE). Performance gains were more pronounced for traits with moderate linkage disequilibrium (LD; r 2 = 0.21-0.36) and high heritability ( h 2 > 0.66), such as those in maize, rapeseed, and soybean. For cotton traits characterized by high LD ( r 2 = 0.74) and lower heritability ( h 2 < 0.50), CLCNet maintained robust performance without degradation. Overall, these results demonstrate that CLCNet is an effective framework for improving genomic prediction accuracy and holds strong potential for practical applications in plant breeding. Short abstract CLCNet is a novel deep learning framework for genomic prediction that integrates contrastive learning with chromosome-aware feature selection. By jointly modeling inter-individual genotype–phenotype variation and chromosomal genomic structure, CLCNet improves prediction accuracy under high-dimensional, low-sample-size conditions. Across four crop species and ten agronomic traits, CLCNet consistently outperformed classical statistical, machine learning, and existing deep learning models. The framework also identified biologically relevant SNPs and candidate genes, demonstrating its potential for practical applications in genomic selection and computational plant breeding. Key points We propose CLCNet, a multi-task deep learning framework that integrates contrastive learning with chromosome-aware feature selection for genomic prediction, under high-dimensional, low-sample-size conditions. The chromosome-aware module explicitly exploits genomic structural information to select representative and informative SNPs across chromosomes. Contrastive learning improves model robustness by stabilizing representation learning and reducing the influence of random effects across samples. By complementing GWAS analyses, CLCNet provides additional insights into genotype–phenotype relationships with potential relevance for gene discovery. Biographical Note Jiangwei Huang is a PhD candidate at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. His research focuses on genomic prediction, deep learning, and computational plant breeding. Zhihan Yang is a PhD candidate at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. Her research interests include genomic prediction and bioinformatics. Rongcheng Han is an associate professor at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. His research focuses on bioinformatics and plant phenomics. Yuqiang Jiang is a professor at the Institute of Genetics and Developmental Biology, Chinese Academy of Sciences. His research interests include plant genomics, plant phenomics and genetic improvement. Organization description The Institute of Genetics and Developmental Biology, Chinese Academy of Sciences, is a leading research institute focusing on genetics, genomics, molecular breeding, bioinformatics, and systems biology in plants and animals.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.357
Threshold uncertainty score0.967

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.010
GPT teacher head0.230
Teacher spread0.219 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueBriefings in BioinformaticsSame topicGenetic Mapping and Diversity in Plants and AnimalsFrench-language works237,207