MétaCan
Menu
Retour à la cohorte
Enregistrement W1789442286 · doi:10.1111/j.1365-2044.2012.07133.x

Lies, damn lies, and statistics*

2012· letter· en· W1789442286 sur OpenAlexaboutno aff
Steve Yentis

Notice bibliographique

RevueAnaesthesia · 2012
Typeletter
Langueen
DomaineMedicine
ThématiqueEthics in Clinical Research
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésMedicineInstitutionLibrary scienceSociologySocial scienceComputer science

Résumé

récupéré en direct d'OpenAlex

In April 2010, this journal carried an editorial entitled ‘Fraud or flawed: adverse impact of fabricated or poor quality research’ [1]. In it, Moore et al. described the difficulty in telling when trial data were fabricated, and how systematic reviews could help by demonstrating aberrations that appeared when data from different studies were presented together. Moore et al. used data from one researcher in Japan to illustrate their point, referring to a previous comment on this researcher’s data that was published 10 years earlier, with the beautifully understated title ‘Reported data on granisetron and postoperative nausea and vomiting by Fujii et al. are incredibly nice!’ [2]. The Journal received a number of responses to Moore et al.’s editorial, amongst them one from a reader who bemoaned the fact that the evidence base remained distorted by that researcher’s work, and challenging anaesthetic journal editors to do something about it. This all happened just a few months after an unprecedented international collaboration between anaesthetic journal editors and publishers, that had led to retraction of almost 90 published papers (including six from this journal) by Joachim Boldt [3]. With this in mind, I counter-challenged the correspondent to perform an analysis of Fujii’s work, more in-depth than that of Kranke etal.’s. If suggestive that Fujii’s data should be removed from the evidence base, this analysis might be used by the same group of editors to confront Fujii and/or his institution(s). So as the correspondent started his work, his letter remained unpublished, and the international group of editors started to discuss how the case might be handled. That correspondent was John Carlisle, and in this issue of the Journal he presents his analysis, some 19 months, 18 reiterations, two consultations with the Committee on Publication Ethics (COPE) and three statisticians later [4]. This is the piece of work that was going to be presented as evidence to Fujii’s institutions – but in the event, it was never used for this purpose, for as we reached the final stages of the manuscript’s preparation, there were rapid developments in Japan. In response to questions raised over a completely separate and coincidental concern that arose in late 2011, Fujii's institution (Toho University, Tokyo) set up an investigation that led to the dismissal of Fujii in February of this year for lack of ethical approval (see Anaesthesia homepage: http://onlinelibrary.wiley.com/journal/10.1111/%28ISSN%291365-2044). Carlisle’s analysis [4] is an extraordinary piece of work. Necessarily based on complicated calculations, its message is nevertheless easily accessible to the reader through its every graph and p value, that clearly show the degree to which Fujii’s data deviate from what would be expected. In an accompanying editorial [5], Pandit provides an explanation of the methods used and their context. Every submission to Anaesthesia is screened using specific software for overlap with other published work before it is reviewed, to detect plagiarism [6]; could Carlisle’s methods similarly form the basis of a screening method for detecting fabrication? I have described in these pages not so long ago the murky world of research misconduct [6]. One question I didn’t cover in any depth is whose responsibility it is to look into suspect cases. This is explored by Wager, in another accompanying editorial [7]. As she points out, institutions don’t always do as they should when approached by editors with concerns – though the institutions in both the Boldt case and the Fujii case should be congratulated for their investigations. As Wager describes, we all have responsibility: from readers through editors, reviewers and publishers to co-investigators, institutions and even governments. From an editorial point of view, being able to access COPE for advice and support, as well as other groups such as the World Association of Medical Editors (see http://www.wame.org/), is enormously useful, and such organisations have a valuable role in keeping up the pressure on all parties, and at all levels, reminding them of their obligations. Medicine – and, it seems, anaesthesia in particular – has seen a flurry of research misconduct scandals recently, with potentially serious implications for the wellbeing of patients who might have received treatments that were ineffective or harmful, and/or been denied those that were beneficial. A further, and possibly longer-lasting, harm is the damage to public trust that results from such cases and the publicity around them. Our specialty is still reeling from the Reuben [8] and Boldt cases [3]; the last thing we needed was another, even bigger, case. We can only hope that the rule that bad things come in threes holds true. Such cases are deeply troubling, raising difficult questions about the framework within which research is done, the motivation and support of those who are prepared to risk so much personally and for their patients, and the reliability of the evidence upon which we rely. However, a lot has changed for the better, since Kranke et al.’s letter a decade ago: there is better awareness of the problem, stronger resolve to stamp out misconduct, and better tools to detect wrongdoing. Carlisle deserves our gratitude for contributing to the last. I am grateful to the authors of these articles and to the Editorial Board, Editorial Team, Publishers, COPE and my fellow Editors-in-Chief of other journals (in particular, Steve Shafer of Anesthesia and Analgesia and Don Miller of Canadian Journal of Anesthesia), for their support. I finished my second and final term as a member of COPE’s Council in March 2012.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,073
score de la tête « metaresearch » (Gemma)0,433
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Intégrité de la recherche
Catégories consensuellesaucune
DomaineSignal candidat: Méthodes · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,993
Score d'incertitude au seuil0,386

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0730,433
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0070,007
Études des sciences et des technologies0,0030,012
Communication savante0,0100,011
Science ouverte0,0020,003
Intégrité de la recherche0,0070,020
Charge utile insuffisante (le modèle a refusé de juger)0,0080,004

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,299
Tête enseignante GPT0,488
Écart entre enseignants0,189 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
DomaineMéthodes
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations26
Publié2012
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAnaesthesiaMême sujetEthics in Clinical ResearchTravaux en français237 207