MétaCan
Menu
Back to cohort
Record W3169848230 · doi:10.1016/j.obhdp.2021.02.003

Same data, different conclusions: Radical dispersion in empirical results when independent analysts operationalize and test the same hypothesis

2021· article· en· W3169848230 on OpenAlexaff
Martin Schweinsberg, Michael B. Feldman, Nicola Staub, Olmo R. van den Akker, Robbie C. M. van Aert, Marcel A. L. M. van Assen, Yang Liu, Tim Althoff, Jeffrey Heer, Alex Kale, Zainab Mohamed, Hashem Amireh, Vaishali Venkatesh Prasad, Abraham Bernstein, Emily V. Robinson, Kaisa Snellman, S. Amy Sommer, Sarah M. G. Otner, David Robinson, Nikhil Madan, Raphael Silberzahn, Pavel Goldstein, Warren Tierney, Toshio Murase, Benjamin Mandl, Domenico Viganola, Carolin Strobl, Catherine Schaumans, Stijn Kelchtermans, Chan Naseeb, S. Mason Garrison, Tal Yarkoni, C.S. Richard Chan, Prestone Adie, Paulius Alaburda, Casper J. Albers, Sara Alspaugh, Jeff Alstott, Andrew A. Nelson, Eduardo Ariño de la Rubia, Arzi Adbi, Štěpán Bahník, Jason Min Baik, Laura Winther Balling, Sachin Banker, David A. A. Baranger, Dale J. Barr, Brenda A. Barros-Rivera, Matt Bauer, Blaise Manga Enuh, Lisa Boelen, Katerina Bohle Carbonell, Robert A. Briers, Oliver Burkhard, Miguel-Angel Canela, Laura Castrillo, Timothy Catlett, Olivia Chen, Michael Clark, Brent Cohn, Alex Coppock, Natàlia Cugueró-Escofet, Paul Curran, Wilson Cyrus-Lai, David Dai, Giulio Valentino Dalla Riva, Henrik Danielsson, Rosaria de F.S.M. Russo, Niko de Silva, Curdin Derungs, Frank Dondelinger, Carolina Duarte de Souza, Blessing Dube, Marina Dubova, Ben Mark Dunn, Peter A. Edelsbrunner, Sara Finley, Nick C. Fox, Timo Gnambs, Yuanyuan Gong, Erin Grand, Brandon Greenawalt, Han Dan, Paul H. P. Hanel, Antony B. Hong, David D. Hood, Justin Hsueh, Lilian Huang, Kent Ngan‐Cheung Hui, Keith A. Hultman, Azka Javaid, Lily J. Jiang, Jonathan Jong, Jash Kamdar, David Kane, Gregor Kappler, Erikson Kaszubowski, Christopher Kavanagh, Madian Khabsa, Bennett Kleinberg, Jens Kouros, Heather Krause, Angelos-Miltiadis Krypotos, Dejan Lavbič, Rui Ling Lee, Timothy Leffel, Wei Yang Lim, Silvia Liverani, Bianca Loh, Dorte Lønsmann, Jia Wei Low, Alton Lu, Kyle MacDonald, Christopher R. Madan, Lasse Hjorth Madsen, Christina Maimone, Alexandra Mangold, Adrienne Marshall, Helena Ester Matskewich, Kimia Mavon, Katherine L. McLain, Amelia McNamara, Mhairi McNeill, Ulf K. Mertens, David I. Miller, Ben Moore, Andrew Moore, Eric Nantz, Ziauddin Nasrullah, Valentina Nejković, Colleen S. Nell, Gustav Nilsonne, Rory Nolan, Christopher E. O'Brien, Patrick O’Neill, Kieran J. O’Shea, Toto Olita, Jahna Otterbacher, Diana Palsetia, Bianca Pereira, Ivan Pozdniakov, John Protzko, Jean-Nicolas Reyt, Travis Riddle, Amal Ridhwan Omar Ali, Ivan Ropovik, Joshua M. Rosenberg, Stéphane Rothen, Michael Schulte‐Mecklenbeck, Nirek Sharma, Gordon Shotwell, Martin Skarzynski, William Stedden, Victoria Stodden, Martin A. Stoffel, Scott Stoltzman, Subashini Subbaiah, Rachael Tatman, Paul H. Thibodeau, Sabina Tomkins, Ana Valdivia, Gerrieke B. Druijff-van de Woestijne, Laura Ferlin Viana, Florence Villesèche, W. Duncan Wadsworth, Florian Wanders, Krista Watts, Jason D. Wells, Christopher E. Whelpley, Andy Won, Lawrence Wu, Arthur Yip, Casey Youngflesh, Ju‐Chi Yu, Arash Zandian, Leilei Zhang, Chava Zibman, Eric Luis Uhlmann

Bibliographic record

VenueOrganizational Behavior and Human Decision Processes · 2021
Typearticle
Languageen
FieldDecision Sciences
Topicscientometrics and bibliometrics research
Canadian institutionsDalhousie UniversityMcGill UniversityUniversity of TorontoSt. Michael's Hospital
FundersInstitut Européen d'Administration des AffairesSchweizerischer Nationalfonds zur Förderung der Wissenschaftlichen ForschungNational Science Foundation
KeywordsOperationalizationPsychologyTest (biology)Empirical researchEconometricsDispersion (optics)Social psychologyStatisticsEpistemologyEconomicsMathematicsPhilosophy

Abstract

fetched live from OpenAlex

In this crowdsourced initiative, independent analysts used the same dataset to test two hypotheses regarding the effects of scientists’ gender and professional status on verbosity during group meetings. Not only the analytic approach but also the operationalizations of key variables were left unconstrained and up to individual analysts. For instance, analysts could choose to operationalize status as job title, institutional ranking, citation counts, or some combination. To maximize transparency regarding the process by which analytic choices are made, the analysts used a platform we developed called DataExplained to justify both preferred and rejected analytic paths in real time. Analyses lacking sufficient detail, reproducible code, or with statistical errors were excluded, resulting in 29 analyses in the final sample. Researchers reported radically different analyses and dispersed empirical outcomes, in a number of cases obtaining significant effects in opposite directions for the same research question. A Boba multiverse analysis demonstrates that decisions about how to operationalize variables explain variability in outcomes above and beyond statistical choices (e.g., covariates). Subjective researcher decisions play a critical role in driving the reported empirical results, underscoring the need for open data, systematic robustness checks, and transparency regarding both analytic paths taken and not taken. Implications for organizations and leaders, whose decision making relies in part on scientific findings, consulting reports, and internal analyses by data scientists, are discussed.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.274
metaresearch head score (Gemma)0.571
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.726
Threshold uncertainty score0.896

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2740.571
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0090.007
Science and technology studies0.0040.014
Scholarly communication0.0140.010
Open science0.0040.010
Research integrity0.0030.006
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.402
GPT teacher head0.509
Teacher spread0.107 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations136
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueOrganizational Behavior and Human Decision ProcessesSame topicscientometrics and bibliometrics researchFrench-language works237,207