MétaCan
Menu
Back to cohort

BALANCING TYPE I AND TYPE II ERROR: RESPONSE TO GORMAN 2008

2008· article· en· W1525230572 on OpenAlexaff
Kathryn Graham

Bibliographic record

VenueAddiction · 2008
Typearticle
Languageen
FieldEconomics, Econometrics and Finance
TopicEconomic Policies and Impacts
Canadian institutionsCentre for Addiction and Mental Health
Fundersnot available
KeywordsRigourInterpretation (philosophy)Point (geometry)PsychologyConventionEpistemologyComputer scienceManagement scienceEngineering ethicsSocial psychologySociologyMathematicsSocial scienceEngineering

Abstract

fetched live from OpenAlex

If the title of my commentary implied that practical knowledge can be gained only at the expense of scientific rigour, this was not my intention. My point was that as researchers and evaluators we should be concerned with both application and rigour, that is, ‘how can research increase practical knowledge while still maintaining high standards of scientific rigour?’[1, p. 414]. I agree with Gorman [2,3] that slanting the interpretation to support a desired outcome is inappropriate if not unethical. However, use of multiple outcomes and considering P values other than 0.05 does not necessarily lead to opportunistic and inaccurate presentations of findings; in fact, rigid adherence to a certain definition of ‘science’ may actually contribute to selective reporting and interpretation [4]. On the contrary, looking beyond the arbitrary criterion of P < 0.05 is critical to making the most accurate assessment of results and seeking the most promising directions for future research and practice. As argued by Rothman and Greenland [5, p. 187], ‘Decisions are inevitably based on results from a collection of studies, and proper combination of the information from the studies requires more than just a classification of each study into ‘significant’ or ‘not significant.’ Thus, degradation of information about an effect into a simple dichotomy is counterproductive, even for decision making, and can be misleading’. The point cannot be made too strongly that the convention of P < 0.05 is merely a mechanism for balancing one type of error against another [6]; it does not define scientific rigour. For practical application, the appropriate balance of these two types of errors depends on the risks. For example, if there is evidence at the P < 0.15 criterion that a product is seriously harmful; this might be considered sufficient to ban the product from the market [4]. On the other hand, more stringent criteria (e.g. P < 0.001) along with a preponderance of evidence would be needed before undertaking massive policy change. Perhaps, consideration of the same pattern of results as those found by Toomey et al.[7] for a different kind of research would help to clarify the issues. Suppose that a randomized control trial of a treatment for cancer found an immediate effect on symptoms that was just short of statistical significance (P = 0.06) but was not sustained. If symptoms were severe and no alternative treatments were available, this might be sufficient evidence to implement the treatment on a trial basis for some cases. In addition, it would likely be deemed warranted to conduct follow-up research to improve power or measurement sensitivity by: increasing the sample size; increasing the dosage or duration of treatment; measuring multiple clinical outcomes to explore effects on different symptoms or on survival time; measuring multiple biological outcomes to better understand treatment mechanisms. In the treatment of serious diseases such as cancer, it is easy to recognize the importance of measuring multiple outcomes and, under some circumstances, adopting promising approaches that do not reach the arbitrary criterion of P < 0.05. The same principle of balancing need for practical knowledge with strength of evidence applies to policy. As noted by Single in contrasting the scientific perspective with that of policy makers, ‘Science aims at the progressive development of an immutable body of knowledge. Policy, on the other hand, is generally concerned with what can be done [and] accomplished now, with limited knowledge, time and resources’[8, p. 1105]. Thus, it is important to recognize that a decision not to recommend a program is still a decision. That is, in practical terms, failing to reject the null hypothesis is the same thing as accepting it. While Gorman argues correctly that it is fallacious to cherry-pick positive outcomes to support a particular point of view, he fails to recognize that it is equally fallacious to interpret the lack of a statistically significant difference as implying no true difference unless the design and measurement of the study were sufficiently sensitive to maximize the chances of correctly identifying a true difference [9]. As I argued in my commentary, one problem in concluding that there was no effect based on the Toomey et al. study was that the outcome involved a single dichotomous measure (served/not served) which may not have been sufficiently sensitive to provide an adequate test of the null hypothesis [7]. Another problem in accepting the null hypothesis when effects do not meet the 0.05 criterion for statistical significance is that many evaluations tend to be grossly underpowered. For example, an analysis [10] of meta-analyses of evaluations of prevention and service programs found an average of 55% of individual studies concluded that the program under study was ineffective when a meta-analysis across the individual studies demonstrated a positive effect. In sum, the ‘critical-rational orientation to hypothesis testing’ recommended by Gorman is no guarantee of scientific rigour if research power and sensitivity are not adequately addressed. Finally, Gorman suggests that considering a decision criterion other than P < 0.05 will result in a waste of scarce resources. This argument assumes that a strict criterion of P < 0.05 will prevent the waste of resources on ‘bad’ programs. In fact, due to a strong demand for immediate community action, time and money is often spent on programs that are completely unevaluated which means that some of these programs will have unintended negative consequences. For example, breathalysers were implemented in some bar settings in Australia at one time with the idea that patrons could self-test and avoid driving if they were over the legal limit. Instead, patrons used the breathalysers to see who could blow highest [11]. To conclude, given the immediate need for knowledge, I maintain, as in my original commentary, that interpretation and application of research findings should involve weighing the preponderance of evidence regarding the risks, benefits and costs of a particular policy or program against the risks, benefits and costs of current alternatives, including inaction. None.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.108
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.036
GPT teacher head0.232
Teacher spread0.195 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2008
Admission routes1
Has abstractyes

Explore more

Same venueAddictionSame topicEconomic Policies and ImpactsFrench-language works237,207