The Utility of Length of Mining Service and Latency in Predicting Silicosis among Claimants to a Compensation Trust
Bibliographic record
Abstract
In the wake of a large burden of silicosis and tuberculosis among ex-miners from the South African gold mining industry, several programmes have been engaged in examining and compensating those at risk of these diseases. Availability of a database from one such programme, the Q(h)ubeka Trust, provided an opportunity to examine the accuracy of length of service in predicting compensable silicosis, and the concordance between self-reported employment and that officially recorded. Compensable silicosis was determined by expert panels, with ILO profusion ≥1/0 as the threshold for compensability. Age, officially recorded and self-reported years of service, and years since first and last service of 3146 claimants for compensable silicosis were analysed. Self-reported and recorded service were moderately correlated (R = 0.66, 95% confidence interval 0.64−0.68), with a Bland−Altman plot showing no systematic bias. There was reasonably high agreement with 75% of the differences being less than two years. Logistic regression and receiver operating characteristic curve analysis were used to test prediction of compensable silicosis. There was little predictive difference between length of service on its own and a model adjusting for length of service, age, and years since last exposure. Predictive accuracy was moderate, with significant potential misclassification. Twenty percent of claimants with compensable silicosis had a length of service <10 years; in almost all these claims, the interval between last exposure and the claim was 10 years or more. In conclusion, self-reported service length in the absence of an official service record could be accepted in claims with compatible clinical findings. Length of service offers, at best, moderate predictive capability for silicosis. Relatively short service compensable silicosis, when combined with at least 10 years since last exposure, was not uncommon.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.020 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".