Response to Toshihide Tsuda, Yumiko Miyano and Eiji Yamamoto [1]
Bibliographic record
Abstract
BACKGROUND: In August 2021, we published in Environmental Health a Toolkit for detecting misused epidemiological methods with the goal of providing an organizational framework for transparently evaluating epidemiological studies, a body of evidence, and resultant conclusions. Tsuda et al., the first group to utilize the Toolkit in a systematic fashion, have offered suggestions for its modification. MAIN BODY: Among the suggested modifications made by Tsuda et al., we agree that rearrangement of Part A of the Toolkit to reflect the sequence of the epidemiological study process would facilitate its usefulness. Expansion or adaptation of the Toolkit to other disciplines would be valuable but would require the input of discipline-specific expertise. We caution against using the sections of the Toolkit to produce a tally or cumulative score, because none of the items are weighted as to importance or impact. Rather, we suggest a visual representation of how a study meets the Toolkit items, such as the heat maps used to present risk of bias criteria for studies included in Cochrane reviews. We suggest that the Toolkit be incorporated in the sub-specialty known as "forensic epidemiology," as well as in graduate training curricula, continuing education programs, and conferences, with the recognition that it is an extension of widely accepted ethics guidelines for epidemiological research. CONCLUSION: We welcome feedback from the research community about ways to strengthen the Toolkit as it is applied to a broader assemblage of research studies and disciplines, contributing to its value as a living tool/instrument. The application of the Toolkit by Tsuda et al. exemplifies the usefulness of this framework for transparently evaluating, in a systematic way, epidemiological research, conclusions relating to causation, and policy decisions. POSTSCRIPT: We note that our Toolkit has, most recently, inspired authors with discipline-specific expertise in the field of Conservation Biology to adapt it for use in the Biological Sciences.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.011 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".