MétaCan
Menu
Back to cohort
Record W4361191099 · doi:10.1093/annweh/wxad017

Prerequisite for Imputing Non-detects among Airborne Samples in OSHA’s IMIS Databank: Prediction of Sample’s Volume

2023· article· en· W4361191099 on OpenAlexaff
Igor Burstyn, Philippe Sarazin, George Luta, Melissa C. Friesen, Laurel Kincl, Jérôme Lavoué

Bibliographic record

VenueAnnals of Work Exposures and Health · 2023
Typearticle
Languageen
FieldMedicine
TopicOccupational exposure and asthma
Canadian institutionsUniversité de MontréalInstitut de recherche Robert-Sauvé en santé et en sécurité du travail
FundersNational Cancer InstituteNational Institutes of Health
KeywordsConfidence intervalStatisticsPrediction intervalEnvironmental scienceCensoring (clinical trials)CartMathematicsGeography

Abstract

fetched live from OpenAlex

INTRODUCTION: The US Integrated Management Information System (IMIS) contains workplace measurements collected by Occupational Safety and Health Administration (OSHA) inspectors. Its use for research is limited by the lack of record of a value for the limit of detection (LOD) associated with non-detected measurements, which should be used to set censoring point in statistical analysis. We aimed to remedy this by developing a predictive model of the volume of air sampled (V) for the non-detected results of airborne measurements, to then estimate the LOD using the instrument detection limit (IDL), as IDL/V. METHODS: We obtained the Chemical Exposure Health Data from OSHA's central laboratory in Salt Lake City that partially overlaps IMIS and contains information on V. We used classification and regression trees (CART) to develop a predictive model of V for all measurements where the two datasets overlapped. The analysis was restricted to 69 chemical agents with at least 100 non-detected measurements, and calculated sampling air flow rates consistent with workplace measurement practices; undefined types of inspections were excluded, leaving 412,201/413,515 records. CART models were fitted on randomly selected 70% of the data using 10-fold cross-validation and validated on the remaining data. A separate CART model was fitted to styrene data. RESULTS: Sampled air volume had a right-skewed distribution with a mean of 357 l, a median (M) of 318, and ranged from 0.040 to 1868 l. There were 173,131 measurements described as non-detects (42% of the data). For the non-detects, the V tended to be greater (M = 378 l) than measurements characterized as either 'short-term' (M = 218 l) or 'long-term' (M = 297 l). The CART models were complex and not easy to interpret, but substance, industry, and year were among the top three most important classifiers. They predicted V well overall (Pearson correlation (r) = 0.73, P < 0.0001; Lin's concordance correlation (rc) = 0.69) and among records captured as non-detects in IMIS (r = 0.66, P < 0.0001l; rc = 0.60). For styrene, CART built on measurements for all agents predicted V among 569 non-detects poorly (r = 0.15; rc = 0.04), but styrene-specific CART predicted it well (r = 0.87, P < 0.0001; rc = 0.86). DISCUSSION: Among the limitations of our work is the fact that samples may have been collected on different workers and processes within each inspection, each with its own V. Furthermore, we lack measurement-level predictors because classifiers were captured at the inspection level. We did not study all substances that may be of interest and did not use the information that substances measured on the same sampling media should have the same V. We must note that CART models tend to over-fit data and their predictions depend on the selected data, as illustrated by contrasting predictions created using all data vs. limited to styrene. CONCLUSIONS: We developed predictive models of sampled air volume that should enable the calculation of LOD for non-detects in IMIS. Our predictions may guide future work on handling non-detects in IMIS, although it is advisable to develop separate predictive models for each substance, industry, and year of interest, while also considering other factors, such as whether the measurement evaluated long-term or short-term exposure.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.035
Threshold uncertainty score0.522

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.122
GPT teacher head0.377
Teacher spread0.255 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueAnnals of Work Exposures and HealthSame topicOccupational exposure and asthmaFrench-language works237,207