Covalent: Interpretable and Discriminative Collective Variables Reveal Ligand-Dependent Switching in Human Cellular Retinol-Binding Protein 2
Bibliographic record
Abstract
Identifying collective variables (CVs) that are both discriminative and interpretable remains a central challenge for enhanced sampling and mechanistic analysis of biomolecular systems. We present Covalent ( Collective variables learnt by a computer ), a supervised machine learning-based CV discovery pipeline that combines a filter-wrapper-substitution feature funnel with an improved, Riemannian-optimized variant of harmonic linear discriminant analysis (GDHLDA) and a post hoc subspace rotation to concentrate pairwise transition information. Applied to unbiased molecular dynamics (MD) trajectories of human cellular retinol-binding protein II (CRBP2) in apo, retinol-bound, and 2-lauroylglycerol (2-LaG)-bound states, Covalent yields linear CVs with clear mechanistic interpretations and better class separability than principal component analysis. The learned CVs highlight gating switches that involve portal loop motion with ligand-induced “locking” of Ser76 and coordinated rearrangements around the cavity (Lys40, Phe57) with reduced Tyr60 flexibility upon binding and indicate a 3–4% decrease in internal void volume consistent with tighter β-barrel packing. Covalent also resolves ligand-specific interaction-state switching: Asp113 toggles mutually exclusive salt bridges with Lys114/Lys132, with 2-LaG strongly biasing the Asp113-Lys114 contact; Glu72 exhibits ligand-dependent hydrogen bonding with Thr60/Gln97. Importantly, when used in well-tempered metadynamics simulations initiated from the apo state, the CVs generate a single dominant free energy basin, with no minima corresponding to the holo conformations, supporting that the holo states are not preorganized in the apo protein but require ligand-induced stabilization. Collectively, these results establish Covalent as a practical route to physically transparent CVs that bridge MD data and mechanism and are readily portable to other problems beyond CRBP2.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".