Automated molecular formula determination by tandem mass spectrometry (MS/MS)
Bibliographic record
Abstract
Automated software was developed to analyze the molecular formula of organic molecules and peptides based on high-resolution MS/MS spectroscopic data. The software was validated with 96 compounds including a few small peptides in the mass range of 138-1569 Da containing the elements carbon, hydrogen, nitrogen and oxygen. A Micromass Waters Q-TOF Ultima Global mass spectrometer was used to measure the molecular masses of precursor and fragment ions. Our software assigned correct molecular formulas for 91 compounds, incorrect molecular formulas for 3 compounds, and no molecular formula for 2 compounds. The obtained 95% success rate indicates high reliability of the software. The mass accuracy of the precursor ion and the fragment ions, which is critical for the success of the analysis, was high, i.e. the accuracy and the precision of 850 data were 0.0012 Da and 0.0016 Da, respectively. For the precursor and fragment ions below 500 Da, 60% and 90% of the data showed accuracy within < or = 0.001 Da and < or = 0.002 Da, respectively. The precursor and fragment ions above 500 Da showed slightly lower accuracy, i.e. 40% and 70% of them showed accuracy within < or = 0.001 Da and < or = 0.002 Da, respectively. The molecular formulas of the precursor and the fragments were further used to analyze possible mass spectrometric fragmentation pathways, which would be a powerful tool in structural analysis and identification of small molecules. The method is valuable in the rapid screening and identification of small molecules such as the dereplication of natural products, characterization of drug metabolites, and identification of small peptide fragments in proteomics. The analysis was also extended to compounds that contain a chlorine or bromine atom.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".