MétaCan
Menu
Back to cohort
Record W2982393829 · doi:10.1373/jalm.2018.028373

Reproducible Research and Reports with R

2019· letter· en· W2982393829 on OpenAlexaff
Janet Simons, Daniel T. Holmes

Bibliographic record

VenueThe Journal of Applied Laboratory Medicine · 2019
Typeletter
Languageen
FieldDecision Sciences
TopicScientific Computing and Data Management
Canadian institutionsUniversity of British Columbia HospitalSt. Paul's HospitalUniversity of British Columbia
Fundersnot available
KeywordsComputer scienceSet (abstract data type)Point (geometry)Data scienceQuality (philosophy)Bar chartData miningInformation retrievalProgramming languageStatisticsMathematics

Abstract

fetched live from OpenAlex

We have all had the experience of having performed a laborious calculation in a spreadsheet program only to later be required to redo the analysis because of the availability of additional data, the discovery of an error, or because the analysis is part of a recurring report (e.g., monthly quality indicators). At that point we may have to return and begin the calculation all over, except we may not even remember what we did, or we may inadvertently perform the analysis in a slightly different way each time. Another common issue is that we may have a data set so large that using a spreadsheet program may be impractical—move scroll bar, watch program freeze, go for coffee. Then there's all the help you did not ask for. Has your spreadsheet program ever converted your data points into dates against your will? (1) Another challenge we have all faced is the preparation of an elaborate figure, the complexity of which makes the use of standard software tools unworkable. The problem of retracing your statistical steps is not merely one of inconvenience. There are serious medical consequences to errors attributable to the effects of spreadsheet programs and software operated through a graphical user interface (2). Fundamentally, the issue is one of reproducibility. The opacity of graphical user interface–based statistical analysis and the importance of research transparency and reproducibility have been highlighted by scientific scandals that could have been avoided through a reproducible research paradigm (3). Anytime we manipulate data with mouse moves, a record of what we have actually done (cleansing, outlier removal, statistical methodology) is lost, and if we made a mistake, it will not be traceable. However, if we prepare our analysis in a programming language, we and others can see what we did, provided we have left the original data set undisturbed. Even better, if we could combine the statistical analysis and authorship process, we could produce an entirely reproducible and transparent scientific report. In the past 10 years there have been serious efforts in the statistical and computational sciences to build tools for creating reports and research papers that are themselves a computer program, down to the tables, the figures, the inline quotations of summative statistics, the handling of references and internal cross-references. If written properly, alteration of a single point of raw data will be reflected throughout the paper when the code to generate the paper is rerun. In light of recent developments in so-called “big data” and “data science,” doubtless the reader is aware of the open-source and freely available R statistical programming language, which can be used to perform all manner of analyses in healthcare and laboratory medicine research. However, parallel to the prolific expansion of the language itself and the massive user base has been the development of tools specifically directed at the production of reproducible reports. These programs can read in the raw data, clean it up, format it properly, perform all calculations, and then generate the output in nearly any desired format: static HTML, HTML Dashboards, PDF, Word, Excel, and PowerPoint. The program can then push the output to a local folder, e-mail it, and even send alerts by text message. The manuscript, having calculated this z score, stores it in a variable, denoted z.score. The calculated value, (z = 3.0902), can then be embedded in the document, as it has been in this sentence. As a more elaborate example of an embedded figure built in real time, Fig. 1 shows a rose plot of turnaround time breakdown for stat ward collections over the course of a year. The code to generate this plot and this entire manuscript is provided as an online supplement (see Data Supplement that accompanies the online version of this article at http://www.jalm.org/content/vol4/issue3). Although there are obvious complexities associated with publishing raw data and source code (particularly if there may be yet-undiscovered findings), it is our belief that the death knell has been sounded for the traditional scientific publishing paradigm of presentation of description and findings but without raw data and source code (5) and that open and reproducible publications will become normative in time. RMarkdown and a number of related open-source tools make automated reports and reproducible research possible.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.239
metaresearch head score (Gemma)0.700
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.761
Threshold uncertainty score0.938

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2390.700
Meta-epidemiology (narrow)0.0040.005
Meta-epidemiology (broad)0.0070.007
Bibliometrics0.0140.016
Science and technology studies0.0040.018
Scholarly communication0.0230.012
Open science0.0090.013
Research integrity0.0110.022
Insufficient payload (model declined to judge)0.1310.134

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.199
GPT teacher head0.415
Teacher spread0.217 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainReproducibility
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueThe Journal of Applied Laboratory MedicineSame topicScientific Computing and Data ManagementFrench-language works237,207