Assessing the quality of evidence in studies estimating prevalence of exposure to occupational risk factors: The QoE-SPEO approach applied in the systematic reviews from the WHO/ILO Joint Estimates of the Work-related Burden of Disease and Injury
Notice bibliographique
Résumé
The World Health Organization (WHO) and the International Labour Organization (ILO) have produced the WHO/ILO Joint Estimates of the Work-related Burden of Disease and Injury (WHO/ILO Joint Estimates). For these, systematic reviews of studies estimating the prevalence of exposure to selected occupational risk factors have been conducted to provide input data for estimations of the number of exposed workers. A critical part of systematic review methodology is to assess the quality of evidence across studies. In this article, we present the approach applied in these WHO/ILO systematic reviews for performing such assessments on studies of prevalence of exposure. It is called the Quality of Evidence in Studies estimating Prevalence of Exposure to Occupational risk factors (QoE-SPEO) approach. We describe QoE-SPEO’s development to date, demonstrate its feasibility reporting results from pilot testing and case studies, note its strengths and limitations, and suggest how QoE-SPEO should be tested and developed further. Following a comprehensive literature review, and using expert opinion, selected existing quality of evidence assessment approaches used in environmental and occupational health were reviewed and analysed for their relevance to prevalence studies. Relevant steps and components from the existing approaches were adopted or adapted for QoE-SPEO. New steps and components were developed. We elicited feedback from other systematic review methodologists and exposure scientists and reached consensus on the QoE-SPEO approach. Ten individual experts pilot-tested QoE-SPEO. To assess inter-rater agreement, we counted ratings of expected (actual and non-spurious) heterogeneity and quality of evidence and calculated a raw measure of agreement (Pi) between individual raters and rater teams for the downgrade domains. Pi ranged between 0.00 (no two pilot testers selected the same rating) and 1.00 (all pilot testers selected the same rating). Case studies were conducted of experiences of QoE-SPEO’s use in two WHO/ILO systematic reviews. We found no existing quality of evidence assessment approach for occupational exposure prevalence studies. We identified three relevant, existing approaches for environmental and occupational health studies of the effect of exposures. Assessments using QoE-SPEO comprise three steps: (1) judge the level of expected heterogeneity (defined as non-spurious variability that can be expected in exposure prevalence, within or between individual persons, because exposure may change over space and/or time), (2) assess downgrade domains, and (3) reach a final rating on the quality of evidence. Assessments are conducted using the same five downgrade domains as the Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach: (a) risk of bias, (b) indirectness, (c) inconsistency, (d) imprecision, and (e) publication bias. For downgrade domains (c) and (d), the assessment varies depending on the level of expected heterogeneity. There are no upgrade domains. The QoE-SPEO’s ratings are “very low”, “low”, “moderate”, and “high”. To arrive at a final decision on the overall quality of evidence, the assessor starts at “high” quality of evidence and for each domain downgrades by one or two levels for serious concerns or very serious concerns, respectively. In pilot tests, there was reasonable agreement in ratings for expected heterogeneity; 70% of raters selected the same rating. Inter-rater agreement ranged considerably between downgrade domains, both for individual rater pairs (range Pi: 0.36–1.00) and rater teams (0.20–1.00). Sparse data prevented rigorous assessment of inter-rater agreement in quality of evidence ratings. We present QoE-SPEO as an approach for assessing quality of evidence in prevalence studies of exposure to occupational risk factors. It has been developed to its current version (as presented here), has undergone pilot testing, and was applied in the systematic reviews for the WHO/ILO Joint Estimates. While the approach requires further testing and development, it makes steps towards filling an identified gap, and progress made so far can be used to inform future work in this area.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,022 | 0,012 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».