0381 Adjustment for multiple comparisons in a job and industry-title analysis of a case-control study of prostate cancer
Bibliographic record
Abstract
Objectives To evaluate the impacts of empirical Bayes (EB) and semi-Bayes (SB) adjustment to account for multiple testing in a hypothesis-generating study of prostate cancer (PCa) risk by occupation and industry. Method The study population comprises 1937 PCa cases and 1995 population controls aged 40–75 years, all residing in Montreal. Odds ratios (OR) and 95% confidence intervals (CI) of PCa risk for ever employment in an occupation and industry were estimated using unconditional logistic regression models adjusted for age, ancestry, and family history of PCa. EB and SB adjustment was applied to the estimates, with prior variances of 0.15, 0.25 and 0.35 selected for SB. Occupation and industry effects were considered mutually exchangeable, with the risk estimates shrunk towards their respective global mean. Results 5 of the 89 occupations and 3 of the 63 industries had a significantly elevated PCa risk prior to EB/SB adjustment, compared to an expected 2 and 1.5 categories due to random chance. The only positive association remaining significant following EB was for subjects ever employed in government (OR=1.4, 95% CI 1.1–1.5). The remaining elevated PCa risks with SB were found for employment in social science occupations (OR=1.5, 95% CI 1.1–2.0) and for forestry workers (OR=1.7, 95% CI 1.1–2.6), in addition to government (OR=1.4, 95% CI 1.1–1.7). The choice of prior variance had a negligible impact on the estimates. Conclusions The use of EB and SB reduced the number of positive associations compared to the unadjusted estimates. The elevated PCa risk observed for employment in government remained consistent across the adjustment approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.057 | 0.098 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.004 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".