0068 Using workers compensation data to estimate injury patterns in inter-provincial workers
Bibliographic record
Abstract
Background The western Canadian province of Alberta attracts skilled workers from across Canada to work in the oilfields. We investigated whether information from Workers Compensation Board (WCB) claims would provide unbiased estimates on the rate of injury in such migrant workers. Work injuries in Alberta are compensated by the Alberta WCB regardless of province of residence. Methods The Alberta WCB provided claims data with home province, sex, age, industry and time lost from work. Denominator data came from Statistics Canada, linking census and taxation information. We also recruited a cohort of workers in Fort McMurray, the hub city for oil and gas, and followed them for 4 months to record work injuries. Results From Statistics Canada, we had 1,720,716 people working in Alberta in 2012 whose home was Alberta and 10403 whose home was Newfoundland. The overall rate of injury (with no correction possible for days employed) was lower in the migrant workers, after adjustment for age, sex and industry. Within claims, the pattern of time loss differed importantly: those from Newfoundland had a marked deficit in claims with time loss 1-28 days (OR=0.18: 95%CI 0.12-0.27). WCB reporting among the 151 cohort members was lowest among those from out of province or recently settled: overall only 38% of loss time injuries were reported. Those in precarious employment were more likely to self-medicate or quit their job to avoid being labelled with a history of injury. Conclusion Injury risk in inter-provincial workers could not be estimated using only WCB data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".