Automated structural assignment of derivatized complex N‐linked oligosaccharides from tandem mass spectra
Bibliographic record
Abstract
Glycoprotein function is controlled by several biological factors, one of them being the structure of carbohydrate chains (glycans) attached to specific amino acids of the protein backbone. Changes in glycan structures have been shown to modify the secondary and tertiary conformation of glycoproteins, thus their function. Powerful analytical tools are available for the characterization of sugar structures, and recently mass spectrometry (MS) has been increasingly useful for this purpose. Manual interpretation of tandem mass spectrum is possible but tedious. Automated interpretation should speed the analysis and enhance the results obtained. A new computer program for automated interpretation of tandem MS spectra of complex N-linked glycans oligosaccharides from mammals will be described. N-Linked oligosaccharides standards were derivatized with 1-phenyl-3-methyl-5-pyrazolone (PMP) and analyzed by matrix-assisted laser desorption/ionization (MALDI)-tandem MS. Simulated tandem mass spectra of other common glycans were also generated to test the algorithm. The MALDI-MS/MS spectra featured resolved isotopic distributions for the [M + H](+) and fragment ions of oligosaccharides. These isotopic distributions complicated the automated analysis of the spectra and were removed to leave only monoisotopic peaks. An algorithm was written for this purpose, yielding simplified tandem mass spectra. Another algorithm is then used to determine the structure of the oligosaccharide. A score is then given to each structure, depending on agreement with experimental results. The program successfully assigned the true structure in 24 out of the 28 cases (86%) and the true structure was among the three top scoring structures in all cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".