Expanding Per- and Polyfluoroalkyl Substances Coverage in Nontargeted Analysis Using Data-Independent Analysis and IonDecon
Bibliographic record
Abstract
Per- and polyfluoroalkyl substances (PFAS) are widespread, persistent environmental contaminants that have been linked to various health issues. Comprehensive PFAS analysis often relies on ultra-high-performance liquid chromatography coupled with high-resolution mass spectrometry (UHPLC HRMS) and molecular fragmentation (MS/MS). However, the selection and fragmentation of ions for MS/MS analysis using data-dependent analysis results in only the topmost abundant ions being selected. To overcome these limitations, All Ions fragmentation (AIF) can be used alongside data-dependent analysis. In AIF, ions across the entire m / z range are simultaneously fragmented; hence, precursor–fragment relationships are lost, leading to a high false positive rate. We introduce IonDecon, which filters All Ions data to only those fragments correlating with precursor ions. This software can be used to deconvolute any All Ions files and generates an open source DDA formatted file, which can be used in any downstream nontargeted analysis workflow. In a neat solution, annotation of PFAS standards using IonDecon and All Ions had the exact same false positive rate as when using DDA; this suggests accurate annotation using All Ions and IonDecon. Furthermore, deconvoluted All Ions spectra retained the most abundant peaks also observed in DDA, while filtering out much of the artifact peaks. In complex samples, incorporating AIF and IonDecon into workflows can enhance the MS/MS coverage of PFAS (more than tripling the number of annotations in domestic sewage). Deconvolution in complex samples of All Ions data using IonDecon did retain some false fragments (fragments not observed when using ion selection, which were not isotopes or multimers), and therefore DDA and intelligent acquisition methods should still be acquired when possible alongside All Ions to decrease the false positive rate. Increased coverage of PFAS can inform on the development of regulations to address the entire PFAS problem, including both legacy and newly discovered PFAS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.012 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".