MétaCan
Menu
Retour à la cohorte
Enregistrement W1525230572 · doi:10.1111/j.1360-0443.2008.02347.x

BALANCING TYPE I AND TYPE II ERROR: RESPONSE TO GORMAN 2008

2008· article· en· W1525230572 sur OpenAlexaff
Kathryn Graham

Notice bibliographique

RevueAddiction · 2008
Typearticle
Langueen
DomaineEconomics, Econometrics and Finance
ThématiqueEconomic Policies and Impacts
Établissements canadiensCentre for Addiction and Mental Health
Organismes subventionnairesnon disponible
Mots-clésRigourInterpretation (philosophy)Point (geometry)PsychologyConventionEpistemologyComputer scienceManagement scienceEngineering ethicsSocial psychologySociologyMathematicsSocial scienceEngineering

Résumé

récupéré en direct d'OpenAlex

If the title of my commentary implied that practical knowledge can be gained only at the expense of scientific rigour, this was not my intention. My point was that as researchers and evaluators we should be concerned with both application and rigour, that is, ‘how can research increase practical knowledge while still maintaining high standards of scientific rigour?’[1, p. 414]. I agree with Gorman [2,3] that slanting the interpretation to support a desired outcome is inappropriate if not unethical. However, use of multiple outcomes and considering P values other than 0.05 does not necessarily lead to opportunistic and inaccurate presentations of findings; in fact, rigid adherence to a certain definition of ‘science’ may actually contribute to selective reporting and interpretation [4]. On the contrary, looking beyond the arbitrary criterion of P < 0.05 is critical to making the most accurate assessment of results and seeking the most promising directions for future research and practice. As argued by Rothman and Greenland [5, p. 187], ‘Decisions are inevitably based on results from a collection of studies, and proper combination of the information from the studies requires more than just a classification of each study into ‘significant’ or ‘not significant.’ Thus, degradation of information about an effect into a simple dichotomy is counterproductive, even for decision making, and can be misleading’. The point cannot be made too strongly that the convention of P < 0.05 is merely a mechanism for balancing one type of error against another [6]; it does not define scientific rigour. For practical application, the appropriate balance of these two types of errors depends on the risks. For example, if there is evidence at the P < 0.15 criterion that a product is seriously harmful; this might be considered sufficient to ban the product from the market [4]. On the other hand, more stringent criteria (e.g. P < 0.001) along with a preponderance of evidence would be needed before undertaking massive policy change. Perhaps, consideration of the same pattern of results as those found by Toomey et al.[7] for a different kind of research would help to clarify the issues. Suppose that a randomized control trial of a treatment for cancer found an immediate effect on symptoms that was just short of statistical significance (P = 0.06) but was not sustained. If symptoms were severe and no alternative treatments were available, this might be sufficient evidence to implement the treatment on a trial basis for some cases. In addition, it would likely be deemed warranted to conduct follow-up research to improve power or measurement sensitivity by: increasing the sample size; increasing the dosage or duration of treatment; measuring multiple clinical outcomes to explore effects on different symptoms or on survival time; measuring multiple biological outcomes to better understand treatment mechanisms. In the treatment of serious diseases such as cancer, it is easy to recognize the importance of measuring multiple outcomes and, under some circumstances, adopting promising approaches that do not reach the arbitrary criterion of P < 0.05. The same principle of balancing need for practical knowledge with strength of evidence applies to policy. As noted by Single in contrasting the scientific perspective with that of policy makers, ‘Science aims at the progressive development of an immutable body of knowledge. Policy, on the other hand, is generally concerned with what can be done [and] accomplished now, with limited knowledge, time and resources’[8, p. 1105]. Thus, it is important to recognize that a decision not to recommend a program is still a decision. That is, in practical terms, failing to reject the null hypothesis is the same thing as accepting it. While Gorman argues correctly that it is fallacious to cherry-pick positive outcomes to support a particular point of view, he fails to recognize that it is equally fallacious to interpret the lack of a statistically significant difference as implying no true difference unless the design and measurement of the study were sufficiently sensitive to maximize the chances of correctly identifying a true difference [9]. As I argued in my commentary, one problem in concluding that there was no effect based on the Toomey et al. study was that the outcome involved a single dichotomous measure (served/not served) which may not have been sufficiently sensitive to provide an adequate test of the null hypothesis [7]. Another problem in accepting the null hypothesis when effects do not meet the 0.05 criterion for statistical significance is that many evaluations tend to be grossly underpowered. For example, an analysis [10] of meta-analyses of evaluations of prevention and service programs found an average of 55% of individual studies concluded that the program under study was ineffective when a meta-analysis across the individual studies demonstrated a positive effect. In sum, the ‘critical-rational orientation to hypothesis testing’ recommended by Gorman is no guarantee of scientific rigour if research power and sensitivity are not adequately addressed. Finally, Gorman suggests that considering a decision criterion other than P < 0.05 will result in a waste of scarce resources. This argument assumes that a strict criterion of P < 0.05 will prevent the waste of resources on ‘bad’ programs. In fact, due to a strong demand for immediate community action, time and money is often spent on programs that are completely unevaluated which means that some of these programs will have unintended negative consequences. For example, breathalysers were implemented in some bar settings in Australia at one time with the idea that patrons could self-test and avoid driving if they were over the legal limit. Instead, patrons used the breathalysers to see who could blow highest [11]. To conclude, given the immediate need for knowledge, I maintain, as in my original commentary, that interpretation and application of research findings should involve weighing the preponderance of evidence regarding the risks, benefits and costs of a particular policy or program against the risks, benefits and costs of current alternatives, including inaction. None.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesCharge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,108
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,036
Tête enseignante GPT0,232
Écart entre enseignants0,195 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2008
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAddictionMême sujetEconomic Policies and ImpactsTravaux en français237 207