MétaCan
Menu
Back to cohort
Record W2024981049 · doi:10.1093/jnci/djk189

Statisticians Set Sights on Observational Studies

2007· article· en· W2024981049 on OpenAlexaboutno aff
R. S. Tuma

Bibliographic record

VenueJNCI Journal of the National Cancer Institute · 2007
Typearticle
Languageen
FieldMathematics
TopicAdvanced Causal Inference Techniques
Canadian institutionsnot available
Fundersnot available
KeywordsObservational studyChecklistStrengthening the reporting of observational studies in epidemiologySet (abstract data type)Quality (philosophy)Research designClinical study designReliability (semiconductor)Randomized controlled trialComputer scienceMedical physicsMedicineMedical educationPsychologyStatisticsClinical trialMathematicsPathologyCognitive psychologyEpistemology

Abstract

fetched live from OpenAlex

The reliability of results from observational studies has been called into question many times in the recent past, with several analyses showing that well over half of the reported findings are subsequently refuted. In an effort to improve the quality of epidemiological studies and their design, an international group is developing publication guidelines for observational trials, similar to the CONSORT guidelines that were adopted for randomized trials. The proposed guidelines, called STROBE (STrengthening the Reporting of OBservational studies in Epidemiology), will provide a checklist that authors and journal editors can use to ensure that the necessary information has been included in a manuscript. The checklist will include both basic study design information, such as stating the specific objective and prespecified hypotheses, and more complex elements regarding the handling of quantitative variables and statistical methods. “The [current] standards for reporting epidemiological studies are not that high,” said Stuart Pocock, Ph.D., professor of medical statistics at the London School of Hygiene and Tropical Medicine, who has been involved in the STROBE effort. “Previously, there was no document that has had international recognition to say, ‘These are the issues that you need to get right when you are trying to publish an epidemiological study.’” Study authors don't describe with enough detail what they did and why, he said. They also don't report the results in a way that clearly separates their primary intent from subsequent analyses. “By tightening up the reporting, it will also feed back and tighten up how they design their studies in the first place.” The hope is that better trial design will improve other researchers’ ability to reproduce the results. With widely-used methods, Peter Austin, Ph.D., found that Leos were more likely to have gastrointestinal bleeding, while Sagittarians were more likely be hospitalized for a broken arm. But when the proper P value was used, the associations faded away. The STROBE contributors—epidemiologists, journal editors, and medical statisticians—are still working on the final version of the statement but hope to publish it in several journals later this year. The current draft is available for comment at http://www.strobe-statement.org . “I think CONSORT has improved things,” Pocock said. “What made it possible is most of the leading medical journals have bought into CONSORT. That, plus the general scientific community's recognition that standards were needed, has allowed CONSORT to make an impact. We are optimistic that STROBE will go a similar way.” Such improvements are needed. John Ioannidis, M.D., professor and chair of the University of Ioannina School of Medicine in Greece, and colleagues have published a series of papers looking at the quality of the scientific literature, including observational studies. In one analysis, Ioannidis examined what factors contribute to whether a study's findings are reproducible. On the basis of that information, he estimates that only about 20% of adequately powered epidemiology studies aimed at uncovering previously unknown associations and generating new hypotheses are likely to be true, even if they show a statistically significant result. In a separate analysis, his team found that the main findings of only one of six highly cited observational studies could be replicated; the conclusions from the other five were refuted. “The combination of selective reporting and multiplicity of analysis undermines the credibility of epidemiological studies,” Ioannidis said at the annual meeting of the American Association for the Advancement of Science in San Francisco. Selective reporting, also called publication bias, comes about partly because so many epidemiological studies are designed to generate new ideas (exploratory trials), not test specific hypotheses. The authors analyze the data and then decide what to publish, focusing on the statistically significant results—without mentioning other factors that were tested but that failed to show a significant association. This type of data mining can uncover potentially interesting associations, but the authors need to fully report how they discovered the relationship and explicitly state that the results need to be reproduced in a study designed to examine that hypothesis, Pocock said. “The epidemiological literature by itself is a literature that practically has ubiquitous statistically significant findings,” Ioannidis said. In a survey of 389 papers reporting the results of observational studies, his research group found that 88% included at least one statistically significant positive association in the abstract, while only 43% reported a nonsignificant association. This finding is despite the fact that nearly all associations are not expected to be statistically significant. That means there are a lot of nonsignificant associations that are not being reported, despite having been tested. “In the meta-analysis literature, where one pools findings across studies, the problem is called the file drawer problem,” said Peter Austin, Ph.D., a senior scientist at the Institute for Clinical Evaluative Sciences in Toronto. “Researchers try to estimate how many unpublished studies are sitting in people's file drawers—or hard drives, now—that are nulls. The assumption is that there are a bunch of unpublished studies out there.” The problem is that the average reader of scientific literature or the popular press is unlikely to consider such issues when reading a title proclaiming that a certain food is associated with a given disease. The burden is thus on researchers and journal editors to publish studies that do not find a sexy new association as well as those that do. “The journals can have a central role in improving the accuracy and transparency of the reported information and the avoidance of selective reporting biases,” Ioannidis said. “It is often debated whether it is the fault of the journal editors, peer reviewers, or authors. I think this is a pseudodebate. We, as scientists, are the editors, peer reviewers, and authors, wearing different hats on different occasions. So, it is our own responsibility to make things better.” Another factor that often leads to false-positive results is that too many hypotheses are being tested simultaneously. To illustrate the problem, Austin looked for an association between an individual's astrological sign and the likelihood of hospitalization for a particular medical problem. And though the question is biologically implausible, his approach parallels that used in many exploratory studies. First, Austin split his sample of more than 10 million individuals in the Ontario government health databases into a test set and a validation set. Then he looked in the test set for two ailments per astrological sign that occurred significantly more often for individuals born under that sign than those born under the 11 others. When he retested those 24 hypotheses against his validation set, two were statistically significant when the standard P -value cutoff of .05 was used: Leos were significantly more likely to suffer from gastrointestinal bleeding than others ( P = .048), and Sagittarians were more likely to be hospitalized because of a broken arm than others ( P = .0125). “If you keep looking, eventually you will find an association.,” Austin said. But Austin's analysis wasn't quite right. If the researcher wants to maintain a false-positive rate of just 5% (which is what a P value of .05 signifies in a single-hypothesis test), then the number of hypotheses being tested needs to be taken into account when setting the P -value cutoff for significance. When 24 hypotheses are tested, an association would have to have a P value less than .00213 to be statistically significant. Not surprisingly, with that boundary all 24 hypotheses were rejected in the astrological test. Adjusting the significance boundary only partially solves the problem, though. If too many hypotheses are tested, adjusting the P value is likely to wipe out all the associations, even those that have a true biological effect. A better approach is for researchers to look for biological plausibility before the study begins and resources are on the line, Pocock said. “Sharpening the mind rather than doing a post-hoc adjustment of P values is a better way to go,” he said. He hopes that STROBE will encourage such forethought and transparency.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.257
Threshold uncertainty score0.386

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.693
GPT teacher head0.578
Teacher spread0.116 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2007
Admission routes1
Has abstractyes

Explore more

Same venueJNCI Journal of the National Cancer InstituteSame topicAdvanced Causal Inference TechniquesFrench-language works237,207