DnsID in MyCompoundID for Rapid Identification of Dansylated Amine- and Phenol-Containing Metabolites in LC–MS-Based Metabolomics
Bibliographic record
Abstract
High-performance chemical isotope labeling (CIL) liquid chromatography-mass spectrometry (LC-MS) is an enabling technology based on rational design of labeling reagents to target a class of metabolites sharing the same functional group (e.g., all the amine-containing metabolites or the amine submetabolome) to provide concomitant improvements in metabolite separation, detection, and quantification. However, identification of labeled metabolites remains to be an analytical challenge. In this work, we describe a library of labeled standards and a search method for metabolite identification in CIL LC-MS. The current library consists of 273 unique metabolites, mainly amines and phenols that are individually labeled by dansylation (Dns). Some of them produced more than one Dns-derivative (isomers or multiple labeled products), resulting in a total of 315 dansyl compounds in the library. These metabolites cover 42 metabolic pathways, allowing the possibility of probing their changes in metabolomics studies. Each labeled metabolite contains three searchable parameters: molecular ion mass, MS/MS spectrum, and retention time (RT). To overcome RT variations caused by experimental conditions used, we have developed a calibration method to normalize RTs of labeled metabolites using a mixture of RT calibrants. A search program, DnsID, has been developed in www.MyCompoundID.org for automated identification of dansyl labeled metabolites in a sample based on matching one or more of the three parameters with those of the library standards. Using human urine as an example, we illustrate the workflow and analytical performance of this method for metabolite identification. This freely accessible resource is expandable by adding more amine and phenol standards in the future. In addition, the same strategy should be applicable for developing other labeled standards libraries to cover different classes of metabolites for comprehensive metabolomics using CIL LC-MS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".