Automated structural assignment of derivatized complex N‐linked oligosaccharides from tandem mass spectra
Bibliographic record
Abstract
Glycoprotein function is controlled by several biological factors, one of them being the structure of carbohydrate chains (glycans) attached to specific amino acids of the protein backbone. Changes in glycan structures have been shown to modify the secondary and tertiary conformation of glycoproteins, thus their function. Powerful analytical tools are available for the characterization of sugar structures, and recently mass spectrometry (MS) has been increasingly useful for this purpose. Manual interpretation of tandem mass spectrum is possible but tedious. Automated interpretation should speed the analysis and enhance the results obtained. A new computer program for automated interpretation of tandem MS spectra of complex N-linked glycans oligosaccharides from mammals will be described. N-Linked oligosaccharides standards were derivatized with 1-phenyl-3-methyl-5-pyrazolone (PMP) and analyzed by matrix-assisted laser desorption/ionization (MALDI)-tandem MS. Simulated tandem mass spectra of other common glycans were also generated to test the algorithm. The MALDI-MS/MS spectra featured resolved isotopic distributions for the [M + H](+) and fragment ions of oligosaccharides. These isotopic distributions complicated the automated analysis of the spectra and were removed to leave only monoisotopic peaks. An algorithm was written for this purpose, yielding simplified tandem mass spectra. Another algorithm is then used to determine the structure of the oligosaccharide. A score is then given to each structure, depending on agreement with experimental results. The program successfully assigned the true structure in 24 out of the 28 cases (86%) and the true structure was among the three top scoring structures in all cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".