MétaCan
Menu
Back to cohort
Record W4251679160 · doi:10.31234/osf.io/dmhk2

Building a collaborative Psychological Science: Lessons learned from ManyBabies 1

2019· preprint· en· W4251679160 on OpenAlexaff
Krista Byers‐Heinlein, Christina Bergmann, Catherine Davies, Michael C. Frank, J. Kiley Hamlin, Melissa Kline Struhl, Jonathan F. Kominsky, Jessica Elizabeth Kosie, Casey Lew‐Williams, Liquan Liu, Meghan Mastroberardino, Leher Singh, Connor Waddell, Martin Zettersten, Mélanie Söderström

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldPsychology
TopicLanguage Development and Disorders
Canadian institutionsUniversity of British Columbia
Fundersnot available
KeywordsWorkflowScale (ratio)Open scienceField (mathematics)Best practiceCitizen scienceNarrativePreferenceDiversity (politics)Computer scienceData scienceEngineering ethicsPsychologyKnowledge managementPublic relationsManagement sciencePolitical scienceEngineering

Abstract

fetched live from OpenAlex

The field of infancy research faces a difficult challenge: some questions require samples that are simply too large for any one lab to recruit and test. ManyBabies aims to address this problem by forming large-scale collaborations on key theoretical questions in developmental science, while promoting the uptake of Open Science practices. Here, we look back on the first project completed under the ManyBabies umbrella – ManyBabies 1 – which tested the development of infant-directed speech preference. Our goal is to share the lessons learned over the course of the project and to articulate our vision for the role of large-scale collaborations in the field. First, we consider the decisions made in scaling up experimental research for a collaboration involving 100+ researchers and 70+ labs. Next, we discuss successes and challenges over the course of the project, including: protocol design and implementation, data analysis, organizational structures and collaborative workflows, securing funding, and encouraging broad participation in the project. Finally, we discuss the benefits we see both in ongoing ManyBabies projects and in future large-scale collaborations in general, with a particular eye towards developing best practices and increasing growth and diversity in infancy research and psychological science in general. Throughout the paper, we include first-hand narrative experiences, in order to illustrate the perspectives of researchers playing different roles within the project. While this project focused on the unique challenges of infant research, many of the insights we gained can be applied to large-scale collaborations across the broader field of psychology.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.094
metaresearch head score (Gemma)0.118
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: Qualitative
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.906
Threshold uncertainty score0.499

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0940.118
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.002
Science and technology studies0.0240.022
Scholarly communication0.0220.028
Open science0.0060.040
Research integrity0.0070.012
Insufficient payload (model declined to judge)0.0130.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.090
GPT teacher head0.438
Teacher spread0.349 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designQualitative
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations13
Published2019
Admission routes1
Has abstractyes

Explore more

Same topicLanguage Development and DisordersFrench-language works237,207