MétaCan
Menu
← Back to cohort
Record W3211179683 · doi:10.31223/x5qp7g

Reproducibility in subsurface geoscience

2021· preprint· en· W3211179683 on OpenAlexaff
Michael Steventon, C. Jackson, Mark Ireland, Matt G. Hall, Marcus R. Munafò, Kathryn Roberts

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldDecision Sciences
TopicScientific Computing and Data Management
Canadian institutionsAgile Scientific (Canada)
FundersImperial College London
KeywordsReproducibilityGovernment (linguistics)ConfidentialityEarth scienceComputer scienceData scienceGeologyChemistry

Abstract

fetched live from OpenAlex

Reproducibility, the extent to which consistent results are obtained when an experiment or study is repeated, sits at the foundation of science. The aim of this process is to produce robust findings and knowledge, with reproducibility being the screening tool to benchmark how well we are implementing the scientific method. However, the re-examination of results from many disciplines has caused significant concern as to the reproducibility of published findings. This concern is well-founded – our ability to independently reproduce results build trust both within the scientific community, between scientists and the politicians charged with translating research findings into public policy, and the general public. Within geoscience, discussions and practical frameworks for reproducibility are in their infancy, particularly in subsurface geoscience, an area where there are commonly significant uncertainties related to data (e.g. geographical coverage). Given the vital role of subsurface geoscience as part of sustainable development pathways and in achieving Net Zero, such as for carbon capture storage, mining, and natural hazard assessment, there is likely to be an increased scrutiny on the reproducibility of geoscience results. We surveyed 347 Earth scientists from a broad section of academia, government, and industry to understand their experience and knowledge of reproducibility in the subsurface. More than 85% of respondents recognised there is a reproducibility problem in subsurface geoscience, with >90% of respondents viewing conceptual biases as having a major impact on the robustness of their findings and overall quality of their work. Access to data, undocumented methodologies, and confidentiality issues (e.g. use of proprietary data and methods) were identified as major barriers to reproducing published results. Overall, the survey results suggest a need for funding bodies, data providers, research groups, and publishers to build a framework and set of minimum standards for increasing the reproducibility of, and political and public trust in, the results of subsurface studies.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.403
metaresearch head score (Gemma)0.650
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.597
Threshold uncertainty score0.736

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4030.650
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0040.003
Bibliometrics0.0090.012
Science and technology studies0.0070.041
Scholarly communication0.0210.019
Open science0.0060.023
Research integrity0.0060.007
Insufficient payload (model declined to judge)0.0090.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.293
GPT teacher head0.446
Teacher spread0.152 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes1
Has abstractyes

Explore more

Same topicScientific Computing and Data Management→French-language works237,207→