Inaccurate endoscopy: a better explanation for placebo‐associated endoscopic ulcers: authors’ reply
Notice bibliographique
Résumé
Sirs, We thank Drs Graham and Chan1 for their comments on our paper,2 which seem remarkably similar to their letter3 in response to a 2004 paper from Laine et al.4 which reported a 5% placebo ulcer rate. This paper was included in our meta-analysis. It seems that Graham and Chan did not interpret our paper correctly. Although we were interested in endoscopic ulcers in placebo users and explored associated risk factors, we did not use the term nor did we attribute the ulcers to any ‘mysterious placebo effect’ as suggested by Graham and Chan. First, we do not believe that placebo is a cause of peptic ulcer (such as by an allergic reaction). Although many placebo users (∼19%) complained of at least one adverse event,5 this does not mean that the placebo caused a pharmacological or physical effect. The mechanism of any placebo effect is complex and relates to environmental and host factors, patients’ expectation (after informed consent) and psychological responses, etc. In an endoscopic trial, after exclusion of baseline risk factors for ulcer (e.g. H. pylori infection, concomitant aspirin/corticosteroids use etc.), placebo users may experience stress and/or anxiety about a treatment effect (if the blinding is effective). Thus, the possibility of an ulcer being found in placebo takers may be higher than in subjects on ‘no treatment’, although we are not aware that this question has ever been formally studied. We agree that it is not clear whether a placebo can be a cause of an endoscopic ulcer. Therefore, we clearly stated that an endoscopic ulcer in placebo arms may reflect ‘background noise’ in part because of uncontrolled environmental factors. Hence, placebo is considered important in clinical trials to provide a reliable estimate of the treatment effect. Second, we emphasized the plausible role of baseline risk factors rather than placebo itself for ulcer occurrence in placebo users. As we reported, compared with the decade of the Lanza trials (1975 to 1989), recent studies included more patients (n = 83 vs. 833 and 1944 for the past three decades respectively) and more high risk patients (vs. all healthy volunteers in the Lanza trials). There were also other baseline factors, which can partly explain why the incidence of placebo ulcer was 3% in the past decade, but 0% across the Lanza trials. As shown in our results,2 the increase in placebo ulcer incidence is associated with some well-known baseline risk factors, especially a previous history of GI events and, for example, low-dose aspirin or corticosteroids in addition to the placebo. Graham and Chan further argue that the endoscopic ulcer point prevalence is low in healthy individuals (1% of 190 in their study) and ulcer incidence should be low after a short interval study when the baseline endoscopy was normal. However, the study population in RCTs is always different from that of the general population or of community cohorts. The majority of the trials we evaluated had a treatment duration ≥4 weeks (up to 12 weeks), which is not ‘short’ term. The overall prevalence of peptic ulcer in a community population was reported as ∼10% in Norway.6 There is no study which has investigated healthy subjects taking placebo alone and who have undergone repeated endoscopies to assess the real ulcer incidence. Such a study with different intervals would provide an estimate of how soon an endoscopic ulcer might develop after starting placebo and when spontaneous healing might occur. It is important to remember that our analysis included studies in arthritis patients; it would be ideal, but regrettably unethical and not practical, to include ‘untreated control groups’ to define the true placebo ulcer rate from that in the natural history of arthritis patients. Graham and Chan raise two other questions: misdiagnosis of endoscopic ulcer (e.g. erosion classified as ulcer) and the correlation between endoscopic ulcer and clinical outcomes. They suggest that placebo provides an estimate of the error of endoscopy in the assessment of mucosal damage and recommend that experienced endoscopists should be involved and endoscopic video images captured and reviewed. We agree that inaccurate endoscopy might be one explanation for the reported endoscopic ulcers in placebo users, but are unlikely to be the major reason in these RCTs. Experienced endoscopists and video capture are now usual in most recent trials, although the published reports may not always state this. Phase I or II drug development RCTs differ from emergency studies which require ‘on-call endoscopists’ and may lack such a detailed protocol. Rather, a rigorous, standardized protocol is required including the definition of an endoscopic ulcer, outcome assessment and validation; tightly scheduled endoscopy, experienced research trained endoscopists who understand the importance of accuracy and validation, and endoscopic image documentation are all pre-requisites. Interestingly, the endoscopic ulcer was arbitrarily defined by Dr Graham as a circumscribed mucosal break with a diameter of at least ≥5 mm with perceptible depth to minimize confusion with NSAIDs-induced discrete, acute mucosal erosions.7 One could argue that an unskilled endoscopist, as suggested by Graham and Chan, is more likely to miss an ulcer (false negative), while an experienced endoscopist is more likely to diagnose an erosion as an ulcer (false positive). Even with photo documentation, different visual angles of the same lesion might lead to different interpretation. We could not assess the study quality by knowing how many endoscopists were really experienced or how many trials employed additional approaches to avoid technical mistakes, but as Dr Laine commented in respect of his study ‘Although we found no significant difference in ulcers between the aspirin and placebo groups, a significant increase in erosions was seen with low-dose aspirin, suggesting that the study endoscopists were consistent in differentiating between ulcers and lesser lesions’.8 It has been long debated whether endoscopic ulcer can be considered as a surrogate for ulcer complications. A recent systematic review indicates that, based on consistent and plausible findings from disparate populations and designs, endoscopic ulcers are a meaningful surrogate endpoint for clinically significant ulcer complications. Both endoscopic ulcers and ulcer complications pointed to the same direction and to a similar extent in four distinct circumstances, although direct progression from an endoscopic ulcer to an ulcer complication has not been demonstrated.9 Large outcome studies would be needed to establish the power of the surrogacy. We agree that there is a difference between the interpretation of endoscopic ulcers and ulcer complications, that, clearly, not all endoscopic ulcers lead to a GI complication and that an ulcer complication is a more relevant clinical end point than an endoscopic ulcer is. However, endoscopic ulcers have been used in numerous trials as a marker of clinically significant outcomes, such as ulcer bleeding, based on the rationale that to prevent an ulcer effectively will inevitably prevent ulcer bleeding.10 Considering that ‘In general there is a reasonable correlation between endoscopic and clinical outcome studies’, as stated by Dr Chan in 2004,11 the presence of an endoscopic ulcer still has a role in the safety assessment of potentially gastro-toxic agents, especially in populations with a low incidence of GI complications. Endoscopic ulcers might ‘come and go quickly’ within the interval between endoscopies and so we might not be able to capture some uncomplicated ulcers. However, before we have validated prospective evidence, we should not categorize a new found endoscopic ulcer to be ‘only a technical mistake’ or label it as being a ‘not-meaningful, invalidated surrogate for clinically significant GI harm only for marketing purpose’. Drs Graham and Chan have repeated their concern about whether an endoscopic ulcer occurring in the placebo arm of an RCT is caused by placebo or technical error. We suggest that they perform a prospective study comparing placebo with no treatment, employing experienced vs. inexperienced endoscopists. We anticipate their results with great interest. Declaration of personal and funding interests: None.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,018 | 0,139 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,003 | 0,003 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,002 | 0,005 |
| Communication savante | 0,003 | 0,008 |
| Science ouverte | 0,004 | 0,002 |
| Intégrité de la recherche | 0,037 | 0,044 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,007 | 0,005 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».