MétaCan
Menu
Back to cohort
Record W2158489906 · doi:10.1177/1745691614548513

You Cannot Step Into the Same River Twice

2014· article· en· W2158489906 on OpenAlexaff
Blakeley B. McShane, Ulf Böckenholt

Bibliographic record

VenuePerspectives on Psychological Science · 2014
Typearticle
Languageen
FieldPsychology
TopicMental Health Research Topics
Canadian institutionsKellogg's (Canada)
Fundersnot available
KeywordsOperationalizationSample size determinationVariation (astronomy)Replication (statistics)Set (abstract data type)Sample (material)Context (archaeology)Statistical powerVariable (mathematics)Computer scienceStatisticsEconometricsPsychologyCognitive psychologyMathematicsEpistemology

Abstract

fetched live from OpenAlex

Statistical power depends on the size of the effect of interest. However, effect sizes are rarely fixed in psychological research: Study design choices, such as the operationalization of the dependent variable or the treatment manipulation, the social context, the subject pool, or the time of day, typically cause systematic variation in the effect size. Ignoring this between-study variation, as standard power formulae do, results in assessments of power that are too optimistic. Consequently, when researchers attempting replication set sample sizes using these formulae, their studies will be underpowered and will thus fail at a greater than expected rate. We illustrate this with both hypothetical examples and data on several well-studied phenomena in psychology. We provide formulae that account for between-study variation and suggest that researchers set sample sizes with respect to our generally more conservative formulae. Our formulae generalize to settings in which there are multiple effects of interest. We also introduce an easy-to-use website that implements our approach to setting sample sizes. Finally, we conclude with recommendations for quantifying between-study variation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.073
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.164
Threshold uncertainty score0.548

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0140.073
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0020.001
Science and technology studies0.0050.008
Scholarly communication0.0060.011
Open science0.0030.007
Research integrity0.0090.014
Insufficient payload (model declined to judge)0.1640.159

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.077
GPT teacher head0.480
Teacher spread0.402 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations95
Published2014
Admission routes1
Has abstractyes

Explore more

Same venuePerspectives on Psychological ScienceSame topicMental Health Research TopicsFrench-language works237,207