Automatic annotation and visualization tool for mass spectrometry based glycomics
Bibliographic record
Abstract
RATIONALE: With the development of glycomics, a large number of glycan structures have been determined by using mass spectrometry (MS)-based techniques. However, most glycan MS data needs to be manually annotated which is time-consuming, unreliable and inaccurate. METHODS: Herein we report a tool for automatically annotating and browsing N-glycan masses and isotopic distributions. We first constructed a training dataset using the Consortium for Functional Glycomics database, in conjunction with data preprocessing and filtering by composition matching. In addition, we improved a matching glycan isotope abundance algorithm through identifying potential overlap region and constructing an optimization model so that it can deconvolute the overlapped glycan isotopic clusters. RESULTS: In the matching process, if the m/z difference of two detected ions was close to an integer from 1 to 5, the m/z range was considered as a potential overlapped region, from the lower m/z to m/z + 5. It was found that there were more than 20 potential overlap regions in each group of data from CHO sample and human testing sample. Because the training dataset was imbalanced, we combined the Supporting Vector Machines (SVMs) algorithm with different sampling techniques, including Synthetic Minority Over-sampling Technique (SMOTE), to classify all potential candidate compositions. The results demonstrated an average of 26.8% increase in annotation sensitivity through the SMOTE-SVMs algorithm. The source code can be obtained from https://sourceforge.net/projects/glycomaid/. CONCLUSIONS: We have developed a new tool which facilities high-throughput glycomics research and assists mass spectrometrists in the interpretation and annotation of glycan samples. Copyright © 2016 John Wiley & Sons, Ltd.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.014 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".