RF-367 Occupational exposure to agents and substances in the CARTaGENE cohort
Bibliographic record
Abstract
Introduction Occupational exposures are related to occupational diseases burden and increased susceptibility to health issues. The joint assessment of occupational exposure and disease outcome is the key to accelerate breakthroughs in occupational health research. Objectives Occupational data including coding of occupations and exposure assessment from a large population cohort such as CARTaGENE may help the research community in uncovering workplace-related health disparities. Methods CARTaGENE is the largest prospective cohort in Quebec with 43,000 participants recruited among the general population aged 40–69 years at baseline. Approximately 10,000 participants filled out an occupational history questionnaire and data were then coded to create a job titles and industry types database. Then, the CANJEM matrix was applied to assign exposure to 258 chemical agents based on occupations (probability and median dose of exposure). Results The 10,895 CARTaGENE participants reported a total of 21,612 jobs, 45% were held by men and 55% by women. For 1,253 jobs (5.8%) occupation code in the NOC 2011 system was impossible to assign because of lacking information. The majority of jobs were in white collar occupations (18.97%). Among the most prevalent exposures (>10% jobs having probability of exposure >25%) in the cohort were solvents, polycyclic aromatic hydrocarbons, cleaning agents, biocides, engine emissions and aliphatic alcohols. Overall, 18 agents have an overall prevalence greater than 5%, while a further 64 agents have a prevalence greater that 1%. Conclusion Such data is relevant from a public health perspective that uses a population-based approach. CARTaGENE has the advantage to integrate a rich collection of data on each participant such as health questionnaires (diseases, lifestyle), physical measures (blood pressure, spirometry), biochemical data (triglyceride, creatinine), genetic data that could be combined to occupational history data. Ultimately, this public resource available to researchers worldwide allows to carry out further research on specific diseases or exposures conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.010 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".