Influence of data acquisition modes and data analysis approaches on non-targeted analysis of phthalate metabolites in human urine
Bibliographic record
Abstract
Humans are often exposed to phthalates and their alternatives, on account of their widespread use in PVC as plasticizers, which are associated with harmful human effects. While targeted biomonitoring provides quantitative information for exposure assessment, only a small portion of phthalate metabolites has been targeted. This results in a knowledge gap in human exposure to other unknown phthalate compounds and their metabolites. Although the non-targeted analysis (NTA) approach is capable of screening a broad spectrum of chemicals, there is a lack of harmonized workflow in NTA to generate reproducible data within and between different laboratories. The objective of this study was to compare two different NTA data acquisition modes, the data-dependent (DDA) and independent (DIA) acquisition (DDA), as well as two data analysis approaches, based on diagnostic ions and Compound Discoverer software for the prioritization of candidate precursors and identification of unknown compounds in human urine. Liquid chromatography coupled to high-resolution mass spectrometry was used for sample analysis. The combination of three-diagnostic-ion extraction and DDA data acquisition was able to improve data filtering and data analysis for prioritizing phthalate metabolites. With DIA, 25 molecular features were identified in human urine, while 32 molecular features were identified in the same urine samples using DDA data. The number of molecular features identified with level 1 confidence was 11 and 9 using DIA and DDA data, respectively. The study demonstrated that besides sample preparation, the impact of data acquisition must be taken into account when developing a NTA method and a consistent protocol for evaluating such an impact is necessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".