MétaCan
Menu
Back to cohort
Record W4293199897 · doi:10.1111/2041-210x.13961

A call for clean code to effectively communicate science

2022· article· en· W4293199897 on OpenAlexafffund
Alessandro Filazzola, Christopher J. Lortie

Bibliographic record

VenueMethods in Ecology and Evolution · 2022
Typearticle
Languageen
FieldDecision Sciences
TopicScientific Computing and Data Management
Canadian institutionsADD CentreOnex (Canada)York UniversityCentre For Cold Ocean Resources EngineeringUniversity of Toronto
FundersNatural Sciences and Engineering Research Council of CanadaUniversity of Toronto
KeywordsComputer scienceCoding (social sciences)Disk formattingCode reviewData scienceSoftwareBest practiceSource codeSoftware engineeringSoftware developmentStatic program analysisProgramming language

Abstract

fetched live from OpenAlex

Abstract Effective coding is fundamental to the study of biology. Computation underpins most research, and reproducible science can be promoted through clean coding practices. Clean coding is crafting code design, syntax and nomenclature in a manner that maximizes the potential to communicate its intent with other scientists. However, computational biologists are not software engineers, and many of our coding practices have developed ad hoc without formal training, often creating difficult‐to‐read code for others. Hard‐to‐understand code can thus be limiting our efficiency and ability to communicate as scientists with one another. The purpose of this paper is to provide a primer on some of the practices associated with crafting clean code by synthesizing a transformative text in software engineering along with recent articles on coding practices in computational biology. We review past recommendations to provide a series of best practices that transform coding into a human‐accessible form of communication. Three common themes shared in this synthesis are the following: (a) code has value and you are responsible for its organization to enable clear communication , (b) use a formatting style to guide writing code that is easily understandable and consistent and (c) apply abstraction to emphasize important elements and declutter. While many of the provided practices and recommendations were developed with computational biologists in mind, we believe there is wider applicability to any biologist undertaking work in data management or statistical analyses. Clean code is thus a crucial step forward in resolving some of the crisis in reproducibility for science.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.179
metaresearch head score (Gemma)0.466
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.991
Threshold uncertainty score0.948

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1790.466
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0090.007
Science and technology studies0.0080.048
Scholarly communication0.0240.050
Open science0.0090.018
Research integrity0.0110.028
Insufficient payload (model declined to judge)0.0070.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.188
GPT teacher head0.511
Teacher spread0.323 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainReproducibility
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations25
Published2022
Admission routes2
Has abstractyes

Explore more

Same venueMethods in Ecology and EvolutionSame topicScientific Computing and Data ManagementFrench-language works237,207