MétaCan
Menu
Back to cohort
Record W2898532196 · doi:10.1093/annweh/wxy044

Development of and Selected Performance Characteristics of CANJEM, a General Population Job-Exposure Matrix Based on Past Expert Assessments of Exposure

2018· article· en· W2898532196 on OpenAlexafffundabout
Jean‐François Sauvé, Jack Siemiatycki, France Labrèche, Lesley Richardson, Javier Pintos, Marie‐Pierre Sylvestre, Michel Gérin, Denis Bégin, Aude Lacourt, Tracy L Kirkham, Thomas Rémen, Romain Pasquet, Mark S. Goldberg, Marie‐Claude Rousseau, Marie‐Élise Parent, Jérôme Lavoué

Bibliographic record

VenueAnnals of Work Exposures and Health · 2018
Typearticle
Languageen
FieldMedicine
TopicOccupational and environmental lung diseases
Canadian institutionsInstitut National de la Recherche ScientifiqueMcGill University Health CentrePublic Health OntarioUniversity of TorontoUniversité du Québec à MontréalUniversité de MontréalArmand Frappier MuseumInstitut de recherche Robert-Sauvé en santé et en sécurité du travailMcGill UniversityCentre Hospitalier de l’Université de MontréalHEC Montréal
FundersMedical Research CouncilNational Cancer InstituteNational Institutes of HealthMedical Research Council CanadaCancer Research SocietyCanadian Institutes of Health ResearchHealth Canada
KeywordsPopulationControl (management)Job-exposure matrixApplied psychologyComputer sciencePsychologyEnvironmental healthMedicineArtificial intelligence

Abstract

fetched live from OpenAlex

Objectives: We developed a job-exposure matrix called CANJEM using data generated in population-based case-control studies of cancer. This article describes some of the decisions in developing CANJEM, and some of its performance characteristics. Methods: CANJEM is built from exposure information from 31673 jobs held by study subjects included in our past case-control studies. For each job, experts had evaluated the intensity, frequency, and likelihood of exposure to a predefined list of agents based on jobs histories and descriptions of tasks and workplaces. The creation of CANJEM involved a host of decisions regarding the structure of CANJEM, and operational decisions regarding which parameters to present. The goal was to produce an instrument that would provide great flexibility to the user. In addition to describing these decisions, we conducted analyses to assess how well CANJEM covered the range of occupations found in Canada. Results: Even at quite a high level of resolution of the occupation classifications and time periods, over 90% of the recent Canadian working population would be covered by CANJEM. Prevalence of exposure of specific agents in specific occupations ranges from 0% to nearly 100%, thereby providing the user with basic information to discriminate exposed from unexposed workers. Furthermore, among exposed workers there is information that can be used to discriminate those with high exposure from those with low exposure. Conclusions: CANJEM provides good coverage of the Canadian working population and possibly that of several other countries. Available in several occupation classification systems and including 258 agents, CANJEM can be used to support exposure assessment efforts in epidemiology and prevention of occupational diseases.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.029
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.054
Threshold uncertainty score0.108

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.029
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0050.003
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0030.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.048
GPT teacher head0.367
Teacher spread0.319 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations30
Published2018
Admission routes3
Has abstractyes

Explore more

Same venueAnnals of Work Exposures and HealthSame topicOccupational and environmental lung diseasesFrench-language works237,207