O-304 Concordance between the Canadian job-exposure matrix (CANJEM) and expert assessment in occupational exposure assessment among jobs held by women
Bibliographic record
Abstract
Introduction The Canadian job-exposure matrix (CANJEM) is a general population JEM built from expert assessment data of 31,673 jobs held by 8,760 participants from four Montreal case-control studies. Objective To examine the validity of CANJEM for jobs held by women, by comparing exposure assessments using CANJEM and our expert assessment method to a selected list of 69 agents. Methods We compared the exposure estimates for 69 agents within a population-based case-control study of lung cancer assigned by expert assessment to those derived from CANJEM. We linked the job histories of 998 women (3403 jobs) to CANJEM and thereby, derived probability of exposure to each of the 69 selected agents in each job. To create binary exposure variables (exposed/unexposed), we dichotomised probability of exposure using two cutpoints: 25% and 50% (referred to as CANJEM-25% and CANJEM-50%). Using the 3403 jobs as units of observation, we estimated the prevalence of exposure to each selected agent using CANJEM-25% and CANJEM-50%, and using expert assessment. Further, using expert assessment as the gold standard, for each agent, we estimated sensitivity, specificity and Kappa. Results CANJEM-based prevalence estimates correlated well with the prevalences assessed by the experts. Sensitivity, specificity and Kappa varied greatly among agents, and between CANJEM-25% and CANJEM-50% probability of exposure. For some agents such as fabric dust and cooking fumes, the concordance between CANJEM-based and expert-based assessments was high and inspired confidence that CANJEM-based assessments will be adequate; however for many other agents, the concordance was low. We present concordance estimates for 69 agents. Conclusion Exposure concordance measures between CANJEM and expert assessment differed greatly by agents. The results of this study could guide users of CANJEM as to which agents are most likely to provide results that mimic those that would be obtained with expert assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.072 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".