Atom-level Machine Learning of Protein-Glycan Interactions and Cross-chiral Recognition in Glycobiology
Bibliographic record
Abstract
Abstract Cross-chiral recognition in glycobiology is the interactions between biologically conventional proteins and the enantiomers of biological glycans (e.g., L-proteins binding with L-hexoses) from organisms across all kingdoms of life. By symmetry, it also describes the interactions of chirally mirrored proteins with normal D-glycans. Knowledge of cross-chiral recognition is critical to understanding the potential interactions of existing life forms with artificial mirror-life forms, but currently known rules of protein-glycan interaction are insufficient. To build a methodology for learning such interactions, we used machine learned models that can predict binding strength between a set of proteins and glycans represented as graphs of atoms, rather than monosaccharides. Atomic q -gram and Morgan fingerprint (MF) based representation of glycans made it possible to train ML models that predict lectin binding properties of glycans, glycomimetic compounds, and enantiomers of all natural glycans. Critical to this training was merging disparate data—some with relative fluorescence units (RFU) from glycan microarrays and others with K d values from ITC—using a universal “fraction bound” parameter f at a specific lectin concentration. A fully-connected neural network architecture, MCNet takes a MF and concentration (C) as inputs and returns f for 147 lectins. Performance of MCNet is comparable to the GlyNet models, and by proxy to other state-of-the art models that predict strength of protein-glycan interactions. MCNet effectively predicts binding of glycomimetic compounds to Galectins 1, 3, and 7. Breaking from a monosaccharide-based description makes it possible for MCNet to predict cross-chiral recognition. We employed a Liquid Glycan Array to validate some predictions, such as the lack of interactions of L-mannose with D-mannose binding lectins, purified ConA, and DC-SIGN displayed on cells, and weak binding of L-Man to galactose-binding lectins and L-Glc binding by canonical fucose binding lectins. MCNet’s atom-level input makes it possible to agglomerate protein-glycan data from diverse glycans across all kingdoms of life and non-glycan structures (e.g. glycomimetic compounds). The universal fraction bound parameter makes it possible to unify disparate quantitative observations ( K d / IC 50 , RFU, chromatographic retention times, etc.). We believe that such an approach will facilitate a merger of knowledge from diverse glycobiology datasets and predict protein interactions with uncommon/unnatural glycans not attainable from current ML models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".