Identifying cumulative trauma disorders of the upper extremity in workers' compensation databases
Bibliographic record
Abstract
BACKGROUND: Impeding the use of workers' compensation databases for surveillance of cumulative trauma disorder of the upper extremity (CTDUE) is the lack of valid and reliable extraction strategies. METHODS: Using the Z795-96 Coding of Work Injury or Disease Information standard, an algorithm was developed to classify claims as definite, possible, or non-CTDUE. Reliability was assessed with standardized claim reviews. RESULTS: Moderate to substantial agreement (Kappa = 0.48, 95% CI 0.42-0.54, n = 328; weighted Kappa = 0.75, 95% CI 0.70-0.80, n = 328) was demonstrated. The algorithm produced relatively homogeneous groups of definite and non-CTDUE claims but 29.1% of the possible CTDUE claims were categorized as definite CTDUE by claim review. Part of body agreement was almost perfect (Kappa = 0.81-1.00) when determining whether the upper extremity or specific parts of the upper extremity were involved. CONCLUSIONS: The algorithm can be used to estimate the number of CTDUE and extract homogeneous groups of definite and non-CTDUE claims. Furthermore, certain upper extremity part of body codes can be used to target anatomically defined claims.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.058 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.011 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".