MétaCan
Menu
Retour à la cohorte
Enregistrement W2106842304 · doi:10.1093/annhyg/met015

Quality of Evidence Must Guide Risk Assessment of Asbestos, by Lenters, V; Burdorf, A; Vermeulen, R; Stayner, L; Heederik, D

2013· letter· en· W2106842304 sur OpenAlexaff

Notice bibliographique

RevueThe Annals of Occupational Hygiene · 2013
Typeletter
Langueen
DomaineMedicine
ThématiqueOccupational and environmental lung diseases
Établissements canadiensMcGill University
Organismes subventionnairesnon disponible
Mots-clésAsbestosEnvironmental scienceEnvironmental healthMedicineMetallurgyMaterials science

Résumé

récupéré en direct d'OpenAlex

In 2011, Lenters et al. (2011) published a meta-analysis of the relationship between asbestos exposure and consequent lung cancer risk. In 2012, this journal published a critique (Berman and Case, 2012) of the meta-analysis, together with a response (Lenters et al., 2012) by (most of) the original authors. This letter examines some statistical issues raised in this exchange, and suggests that the main finding of the original meta-analysis is less robust than claimed, and should be interpreted cautiously. Lenters et al. performed their calculations in SAS, using a random effects approach and estimating between study variance using restricted maximum likelihood. The present calculations are made on the same basis, but using the R package ‘metafor’ (Viechtbauer, 2010). The central claim of the original meta-analysis was that ‘studies with higher quality asbestos exposure assessment yield higher meta-estimates of the lung cancer risk per unit of exposure’, and the essence of Berman and Case’s critique was that this conclusion was based on estimates that were in many cases not conventionally statistically significant, and were additionally over-reliant on the findings from 1 of the 19 studies examined. Table 1 of Lenters et al. (2011) lists the studies, and this letter uses the numbering given there. Permutation tests showing the sensitivity of the finding of associations of quality criteria with higher KL estimates to the status of Study 3 (satisfying no criteria) and Studies 4 and 9 (satisfying all criteria). Coverage, >30% of work histories covered by exposure measurement data; CE ratio, ratio of cumulative exposures in the highest and lowest exposure categories >50; documentation, sufficient documentation of exposure data; conversion, exposures based on fibre counting, or converted from particle counts using internal evidence; job history, individual job histories sufficiently complete (see Lenters et al., 2011 for full definitions). Permutation tests showing the sensitivity of the finding of associations of quality criteria with higher KL estimates to the status of Study 3 (satisfying no criteria) and Studies 4 and 9 (satisfying all criteria). Coverage, >30% of work histories covered by exposure measurement data; CE ratio, ratio of cumulative exposures in the highest and lowest exposure categories >50; documentation, sufficient documentation of exposure data; conversion, exposures based on fibre counting, or converted from particle counts using internal evidence; job history, individual job histories sufficiently complete (see Lenters et al., 2011 for full definitions). In their response, Lenters et al. (2012) reiterate the three lines of meta-analytical finding on which they base their conclusion: For each of five predefined binary quality criteria, the meta-estimate of the exposure–response coefficient (KL) for studies satisfying the criterion was greater than that for studies that did not; This observation remains true when these estimates were controlled for fibre type (classified as chrysotile versus amphibole and mixed); and As the range of studies considered is successively reduced by excluding studies not satisfying the quality criteria taken in turn (start with all studies, then studies satisfying Criterion A, next those satisfying A and B, and so on), the meta-estimate of the exposure response increases both overall and at almost every step, for a range of exclusion orderings. Additional analyses in their response show that Findings 1 and 2 remain true when excluding the study identified by Berman and Case as influential (Study 4, Hein et al., 2007), and that Finding 3 is robust over additional exclusion orderings. They argue that although a number of these individual findings are not conventionally statistically significant, the consistency of positive finding makes it ‘… unlikely that this pattern was observed due to chance, and propose that quality really does matter’. At first sight, these are persuasive observations. Here are five independently constituted quality markers all of which give higher values of KL when satisfied: thinking in terms of a null hypothesis that there is no underlying relationship, this would have a P-value of 1 in 32 (25). However, the five criteria are not statistically independent. Dependence is induced by the overlapping membership of the quality criteria classes (as Lenters et al. note elsewhere). In particular, Studies 4 and 9 (Sullivan, 2007) are members of all these classes, and Study 3 (McDonald et al., 1984) is not included in any of them. How strong an influence does this commonality of quality class membership have on the observed consistency of positive associations and the occurrence of near monotonic increases with successive exclusion? These questions can be formally tested using permutation tests. We look first at the finding that studies selected as satisfying each individual quality criterion in turn give a higher meta-estimate of KL than the remaining studies. How important to the consistency of this finding across quality criteria is it that Studies 4 and 9 are included in each of them, and Study 3 in none of them? Table 1 shows the results from a series of permutation tests addressing this question. Taking the coverage criterion for example, 7 of the 19 studies satisfied this criterion, and we can assess the importance of, say, Study 4 to the finding that these 7 studies give a higher KL value than the remaining 12 by seeing how often a group of studies including Study 4 and 6 others chosen at random would give a higher KL value than the remaining 12. Out of 10 000 groups chosen in this way, 7293 yielded a higher KL value. In other words, just having Study 4 in a group of 7 means that the odds are better than 2:1 that group will have a higher KL value than the remaining 12 studies. Similarly, out of 10 000 groups of 7, randomly chosen but excluding Study 3 (i.e. always having Study 3 in the remainder group), 7141 yielded a higher KL value. These tests (reported in the first four columns of Table 1) examine each criterion independently. Column five shows the probability of finding that the meta-KL values for the studies satisfying the criteria is larger than those for studies that do not, for all five criteria simultaneously. Here, in order to allow for the pattern of overlapping quality group membership among the 19 studies, we randomly permute the 19 observed sets of criterion group membership (equivalently, think of randomizing the observed KL values among the 19 studies) and see how often these randomized data sets have the property that all five quality criterion groups have higher meta-KL values than their complements. To assess the impact of Studies 3, 4, and 9 on these probabilities, we can remove one or more of these studies from the randomization so that for the un-randomized study(ies), the observed pairing of quality criteria and KL is fixed. Thus, for example, when Study 3 is not randomized, the Study 3 KL value always falls in the groups consisting of studies not meeting the relevant criterion. It is clear that for any division of these studies into two groups, there is a strong tendency for the group not including Study 3 to have a higher KL estimate. For each quality criterion, >70% of its randomized replicates constrained to exclude Study 3 will give a higher estimate of KL, and there is a 38% chance that all five criteria will yield a higher estimate. There is a similar, but weaker tendency with the inclusion of Study 4; the individual probabilities of a higher estimate are ~60%, and the chances that all five will be higher is 14%. Applying both these conditions, the probability of finding a higher estimate for all five criteria is nearly three-quarters (72%). When no constraint is placed on selection, the empirical P-values are close to their theoretical value of 0.5, and the probability that five random selections (of the appropriate sizes) will give higher KL estimates is 0.054, somewhat >1 in 32 (0.031) that would apply for five independent variables due to the pattern of common membership of quality groups in the data. Constraining the selected groups to include Study 9 makes little difference to the probabilities. The probability of 0.054 recorded when no constraints are laid on the selection of studies supports Lenters et al.’s (2012) contention that the occurrence of a higher estimate for KL associated with each of the five quality criteria is a feature of the data worth taking seriously. However, the fact that this probability increases to 0.38 if the selections are constrained to exclude Study 3 suggests that the dominant content of this data feature is that Study 3 fails all the quality criteria and has a low (actually, negative) value of KL, rather than that there is an across-the-board relationship between all five criteria and KL. A similar, though slightly more complicated, permutation test can be constructed to test the significance of the near monotonic increase in KL as studies are successively removed. It is obvious that there must always be a net increase overall because the mean for all studies is 0.13 and that for the group of two studies satisfying all criteria is 0.55. At issue is the extent to which the series of KL values increases at each step. Table 4 of Lenters et al. (2012) shows five different exclusion orderings. Each ordering defines a series of nested groups of decreasing size. Table 2 reproduces the data in this table adding a count of how many of the steps in KL from set to set are positive (np) and how often a sequence of sets of the same sizes always including Study 4 but otherwise randomly chosen gives np or more positive steps. All these probabilities are comfortably >0.5. In other words, the near monotonic increase in meta-KL values is exactly what one should expect given that Study 4 is a member of the final group. Thus, beyond the fact that Study 4 satisfies all five quality criteria and has a high KL value, the near monotonic increase in meta-KL values does not constitute evidence that quality (measured by these five criteria) and KL are positively associated. Permutation test of the significance of the observed degree of monotonicity in the sequences of KL estimates shown in Lenters et al.’s (2012) Table 4, given that the final group always contains Study 4. Permutation test of the significance of the observed degree of monotonicity in the sequences of KL estimates shown in Lenters et al.’s (2012) Table 4, given that the final group always contains Study 4. In summary, the two features of the data that Lenters et al. make most of insupport of their main claim are sensitive to a couple of the studies. The consistently higher estimate of KL in studies satisfying each quality criterion is mainly dependent on Study 3 (rather than Study 4 focused on by Berman and Case), whereas the monotonic trend of KL with successive exclusion on quality grounds is entirely explained by the presence of Study 4 in the highest quality group. This does not mean that there is no tendency for studies with good exposure estimates to produce higher estimates of KL (other things being equal). Indeed, as pointed out by Lenters et al., there are good statistical reasons for expecting that this should be the case. And two of the criteria considered individually (coverage and job histories) do show near significant relationships with KL (P = 0.08 in Table 2 of Lenters et al., 2011). But the evidence that there is a general relationship involving all five criteria depends heavily on Studies 3 and 4. The discussion thus far has assumed that the five chosen quality criteria are valid indicators of individual study quality. This can be questioned. First of all, the CE ratio criterion—that the ratio of cumulative exposure in the highest exposure category should be at least 50 times the level in the lowest exposure category—is not a marker of exposure estimate quality. It indicates power to detect a positive exposure–response slope, but the estimates from lower power studies will not be biased simply because the range of exposures observed is narrow. Such studies are therefore fully qualified for inclusion in a meta-analysis. Next, the other criteria adopted are not infallible indicators of exposure estimation quality. The fundamental issue is whether the exposure estimates used are, in fact, correct, and this could be the case when none of the criteria apply. Furthermore, even when all these criteria do apply, estimates can still be erroneous. This is, most likely, illustrated by the example of the Gustavsson et al. (2002) case–control study (Study 19), which satisfies four of the five quality criteria (or three out of four if CE ratio is discounted), but produces a clearly aberrant exposure–response estimate (nearly 10 times higher than the next highest individual study estimate and >100 times the average for the 19 studies). This does not mean that the quality criterion approach is not a reasonable one; the criteria are all ones which one would prefer to see fulfilled rather than not. Over a large population of studies, the approach would be robust, but over the small number in play here, the fact that these criteria may not reflect true exposure estimation quality at the individual study level means that we should be cautious in interpreting the results as if they did. Lenters et al. draw two wider conclusions from their analysis. First, that risk assessment should be based only on higher quality studies (though how this should be done is not fully specified); and secondly, that their findings ‘cast doubt on assertions that the epidemiological evidence for lung cancer strongly supports the difference in potency difference asbestos fibre types’. There would be wide agreement that study quality—especially the quality of exposure estimates—should be taken account of in reviewing the epidemiological evidence. The difficult question is how this can best be achieved. Berman and Case (2012) defend the approach adopted by Berman and Crump (2008a, b) of widening the confidence limits on estimates from lower quality studies. They argue that, in the absence of clear evidence of bias, this approach will enhance statistical power compared with an approach that excludes some studies altogether. Lenters et al. are not explicit about the implications for risk assessment of their finding of higher KL values with better quality studies, though they seem to imply that the best estimate is the one that results from pooling only those studies satisfying all the adopted quality criteria—in this case, just two studies. There is probably no ideal solution to this difficulty. As pointed out by Lenters et al., the Berman and Crump approach involves some subjective quantification, and ignores the fact that slopes may be biased and uncertain. On the other hand, there are elements of subjectivity in the quality criteria, and taking an estimate from just two studies puts great weight on having correctly identified the best studies. It also assumes that their findings are generalizable. In a context where there is considerable heterogeneity between the different studies and a number of possible explanatory factors for these differences—‘misclassification of the endpoint, average age of first exposure, residual confounding due to smoking, fibre type and differences in distributions of fibre dimensions’ (Lenters et al., 2012)—implicitly ascribing all the difference between these two studies and the 17 others to their (possible) advantage in exposure estimate quality, goes beyond the strength of the available evidence. There are practical and precautionary reasons for setting limits on the assumption that the per fibre risks lie at the higher end of the observed scale. But this does not imply that these estimates are the closest approach to the truth from a scientific point of view. Turning to their statement about the strength of evidence for a difference in potency between fibre types, Lenters et al. (2011) write that their analyses ‘cast doubt on the conclusion that the epidemiological evidence for lung cancer strongly supports a difference in potency for different fibre types’. The truth of this claim depends on what one means by the word ‘strongly’ and how widely one draws one’s definition of epidemiological evidence. However, even restricting attention to the evidence considered by Lenters et al., fibre type is the most highly associated variable with KL (P = 0.06 in Table 2 of Lenters et al., 2011). The association is not highly significant, and may also be sensitive to particular studies, but the evidence for it is at least as strong as that for the quality criteria, and arguably stronger. In this context, it seems somewhat illogical to use the weaker evidence of association with exposure data quality to reach a strong conclusion about restricting the basis of admissible evidence and then cast doubt on the evidence for differences in potency by saying ‘potency differences for predominantly chrysotile versus amphibole exposed cohorts become difficult to ascertain when meta analyses are restricted to studies with fewer exposure assessment limitations’ (Lenters et al., 2011). In truth, the heterogeneity and apparent inconsistency of the evidence based on this topic means that no statistical conclusions can be as clear and robust as we would like them to be on such an important question. Lenters et al. (2011) are surely right that ‘only further research will satisfactorily clarify the controversial issue of fibre-specific potencies and, furthermore, is warranted considering the politically sensitive nature of this question and the widespread public health impact of historic and current asbestos use’.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,283
score de la tête « metaresearch » (Gemma)0,702
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesMétarecherche
DomaineSignal candidat: Méthodes · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,717
Score d'incertitude au seuil0,884

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,2830,702
Méta-épidémiologie (sens strict)0,0020,002
Méta-épidémiologie (sens large)0,0060,007
Bibliométrie0,0110,009
Études des sciences et des technologies0,0030,011
Communication savante0,0250,017
Science ouverte0,0060,007
Intégrité de la recherche0,0150,030
Charge utile insuffisante (le modèle a refusé de juger)0,0060,004

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,231
Tête enseignante GPT0,445
Écart entre enseignants0,213 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.

Devis d'étudeSans objet
DomaineMéthodes
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations2
Publié2013
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueThe Annals of Occupational HygieneMême sujetOccupational and environmental lung diseasesTravaux en français237 207