Impact of breast density on classification of infrared spectroscopy for breath-based breast cancer screening.
Bibliographic record
Abstract
e13558 Background: While mammography is currently the standard of care for breast cancer screening, dense breast tissue can significantly degrade results. Alternatively, infrared spectroscopy analysis of breath offers a highly sensitive method for identifying exhaled volatile organic compounds (VOCs), which may circumvent issues with breast density. Methods: Alveolar breath samples were collected onto Tenax TA sorbent tubes using a SohnoXB™ breath sampler. Using four desorb temperatures (75, 150, 225 and 300°C), absorption spectra were measured by infrared cavity ring-down spectroscopy (IR-CRDS), a technique for measuring absorption coefficients due to VOCs in exhaled breath. Missing values in the absorption spectra were backfilled using interpolation, and the spectrum was min-max normalized and quadratic detrended. First and second derivatives of the preprocessed absorption spectra were used as features for a support vector machine machine-learning model. Features were ranked based on minimum redundancy maximum relevance (mRMR). The top 20 ranked features were selected to limit the potential for overfitting and then optimized. Model performance was validated using non-nested leave-one-out cross-validation (LOOCV) and nested LOOCV to provide optimistic and pessimistic results, respectively. Results: Absorption spectra from 111 participants (71 positive, 40 control) were used. Of the positive subjects, 30 had low- and 31 had high-density breast tissue (measures missing for 10). Model performance is outlined. A subgroup analysis compared model performance for subjects with low- and high-density breast tissue for both non-nested and nested LOOCV models. Breast density data was not captured for control subjects thus they were not included in the subgroup analysis. Fisher’s exact test was performed to assess for significant difference between model performance for those with low- vs high-density breast tissue, resulting in a p-value of 1.00. Conclusions: Our results suggest that the classification of alveolar breath using IR-CRDS is a promising technique for the detection of breast cancer that is independent of breast density. Breath analysis may therefore become a new alternative or corroborative to mammography to support clinical decision-making. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".