Bibliographic record
Abstract
If the title of my commentary implied that practical knowledge can be gained only at the expense of scientific rigour, this was not my intention. My point was that as researchers and evaluators we should be concerned with both application and rigour, that is, ‘how can research increase practical knowledge while still maintaining high standards of scientific rigour?’[1, p. 414]. I agree with Gorman [2,3] that slanting the interpretation to support a desired outcome is inappropriate if not unethical. However, use of multiple outcomes and considering P values other than 0.05 does not necessarily lead to opportunistic and inaccurate presentations of findings; in fact, rigid adherence to a certain definition of ‘science’ may actually contribute to selective reporting and interpretation [4]. On the contrary, looking beyond the arbitrary criterion of P < 0.05 is critical to making the most accurate assessment of results and seeking the most promising directions for future research and practice. As argued by Rothman and Greenland [5, p. 187], ‘Decisions are inevitably based on results from a collection of studies, and proper combination of the information from the studies requires more than just a classification of each study into ‘significant’ or ‘not significant.’ Thus, degradation of information about an effect into a simple dichotomy is counterproductive, even for decision making, and can be misleading’. The point cannot be made too strongly that the convention of P < 0.05 is merely a mechanism for balancing one type of error against another [6]; it does not define scientific rigour. For practical application, the appropriate balance of these two types of errors depends on the risks. For example, if there is evidence at the P < 0.15 criterion that a product is seriously harmful; this might be considered sufficient to ban the product from the market [4]. On the other hand, more stringent criteria (e.g. P < 0.001) along with a preponderance of evidence would be needed before undertaking massive policy change. Perhaps, consideration of the same pattern of results as those found by Toomey et al.[7] for a different kind of research would help to clarify the issues. Suppose that a randomized control trial of a treatment for cancer found an immediate effect on symptoms that was just short of statistical significance (P = 0.06) but was not sustained. If symptoms were severe and no alternative treatments were available, this might be sufficient evidence to implement the treatment on a trial basis for some cases. In addition, it would likely be deemed warranted to conduct follow-up research to improve power or measurement sensitivity by: increasing the sample size; increasing the dosage or duration of treatment; measuring multiple clinical outcomes to explore effects on different symptoms or on survival time; measuring multiple biological outcomes to better understand treatment mechanisms. In the treatment of serious diseases such as cancer, it is easy to recognize the importance of measuring multiple outcomes and, under some circumstances, adopting promising approaches that do not reach the arbitrary criterion of P < 0.05. The same principle of balancing need for practical knowledge with strength of evidence applies to policy. As noted by Single in contrasting the scientific perspective with that of policy makers, ‘Science aims at the progressive development of an immutable body of knowledge. Policy, on the other hand, is generally concerned with what can be done [and] accomplished now, with limited knowledge, time and resources’[8, p. 1105]. Thus, it is important to recognize that a decision not to recommend a program is still a decision. That is, in practical terms, failing to reject the null hypothesis is the same thing as accepting it. While Gorman argues correctly that it is fallacious to cherry-pick positive outcomes to support a particular point of view, he fails to recognize that it is equally fallacious to interpret the lack of a statistically significant difference as implying no true difference unless the design and measurement of the study were sufficiently sensitive to maximize the chances of correctly identifying a true difference [9]. As I argued in my commentary, one problem in concluding that there was no effect based on the Toomey et al. study was that the outcome involved a single dichotomous measure (served/not served) which may not have been sufficiently sensitive to provide an adequate test of the null hypothesis [7]. Another problem in accepting the null hypothesis when effects do not meet the 0.05 criterion for statistical significance is that many evaluations tend to be grossly underpowered. For example, an analysis [10] of meta-analyses of evaluations of prevention and service programs found an average of 55% of individual studies concluded that the program under study was ineffective when a meta-analysis across the individual studies demonstrated a positive effect. In sum, the ‘critical-rational orientation to hypothesis testing’ recommended by Gorman is no guarantee of scientific rigour if research power and sensitivity are not adequately addressed. Finally, Gorman suggests that considering a decision criterion other than P < 0.05 will result in a waste of scarce resources. This argument assumes that a strict criterion of P < 0.05 will prevent the waste of resources on ‘bad’ programs. In fact, due to a strong demand for immediate community action, time and money is often spent on programs that are completely unevaluated which means that some of these programs will have unintended negative consequences. For example, breathalysers were implemented in some bar settings in Australia at one time with the idea that patrons could self-test and avoid driving if they were over the legal limit. Instead, patrons used the breathalysers to see who could blow highest [11]. To conclude, given the immediate need for knowledge, I maintain, as in my original commentary, that interpretation and application of research findings should involve weighing the preponderance of evidence regarding the risks, benefits and costs of a particular policy or program against the risks, benefits and costs of current alternatives, including inaction. None.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".