ECLARE: multi-teacher contrastive learning via ensemble distillation for diagonal integration of single-cell multi-omic data
Bibliographic record
Abstract
Integrating multimodal single-cell data such as scRNA-seq and scATAC-seq is key for decoding gene regulatory networks. Still, integration remains challenging due to issues related to feature harmonization and limited quantity of paired data. To address these challenges, we introduce ECLARE , a novel framework combining multi-teacher ensemble knowledge distillation with contrastive learning for integrating unpaired single-cell multi-omic data. Briefly, ECLARE trains teacher models on paired datasets to guide a student model for aligning unpaired data, leveraging a refined contrastive objective and optimal-transport-based loss for precise cross-modality alignment. In computational benchmarking, experiments demonstrate ECLARE ’s competitive performance in cell pairing accuracy, multimodal integration and biological structure preservation, indicating that multi-teacher knowledge distillation provides an effective means to improve a diagonal integration model beyond its zero-shot capabilities. In biological case studies, we demonstrate ECLARE ’s applicability using unpaired snRNA-seq and snATAC-seq datasets in major depressive disorder (MDD). Firstly, our results revealed transcription factors and target gene combinations differentially regulated in depression with sex- and cell-type specificity. These findings further reveal gene regulatory interactions in excitatory neurons that are highly relevant to MDD neuropathology, such as those involving EGR1, SOX2 , and NR3C1 . Secondly, we show that ECLARE can learn continuous data manifolds useful for deciphering longitudinal biological processes in neurodevelopment and disease, revealing altered neurodevelopmental programs as potential regulators of depression in females, strongly associated with EGR1 target genes. Altogether, we propose ECLARE as a robust solution for diagonal integration of unpaired multimodal single-cell data that enables the study of altered gene regulation in disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".