MétaCan
Menu
Back to cohort
Record W4417146198 · doi:10.3897/biss.9.180327

Rescuing the Past to Prepare for the Future: Environmental Data Rescue as a Key Activity of Open Science

2025· article· en· W4417146198 on OpenAlexafffundabout
Diane S. Srivastava, David A. G. A. Hunt, Sandra A. Binning, Sandra Emry, Ellen K. Bledsoe, Jessica Reemeyer, Jason Pither

Bibliographic record

VenueBiodiversity Information Science and Standards · 2025
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsUniversity of British Columbia, Okanagan CampusUniversity of ReginaUniversité de MontréalCBC (Canada)Kelowna General HospitalUniversity of British Columbia
FundersNatural Sciences and Engineering Research Council of Canada
KeywordsCustodiansEnvironmental dataResource (disambiguation)Open dataCitizen scienceWork (physics)Key (lock)Data access

Abstract

fetched live from OpenAlex

The need for data rescue. Data are dying all around us: when they are stored in inaccessible or fragile repositories, in proprietary formats, on defunct media or without complete metadata. In environmental science, such data loss has high societal and scientific costs. For example, failure to archive ecological data focused on the Exxon Valdez oil spill represents an estimated loss of >100 million USD (Bledsoe et al. 2022). Loss of environmental data can also represent erasure of historic baselines that can never be recollected. The Living Data Project ( LDP ). This award-winning program helps scientists and organizations archive valuable “legacy” datasets, making them permanently open and accessible (Bledsoe et al. 2022). We pair data custodians with graduate students who we train in reproducible data management and provide mentoring by postdoctoral researchers and faculty and assistance by undergraduates. By training Canada’s next generation of scientists, we aim to ensure that data are never lost again. Over the last five years, we have trained 305 students, across 93 internships, rescuing > 2000 years of data. LDP student working groups then analyze archived data to investigate pressing biodiversity questions and develop open access teaching resources. Benefits of data rescue. LDP-rescued datasets, many stretching back into the last century, provide missing historical baselines that enable the detection of environmental change – such as change in lake ice formation and melt over the last 150 years. Rescue of historical datasets, like surveys of urban birds or plant communities collected five decades ago, sets the stage for contemporary re-surveys. Combining rescued data with contemporary data can also yield discoveries such as legacy effects of pesticides on stream health (Sugden et al. 2025). LDP rescue of classic studies in ecological and evolutionary theory (Darwin’s Finches, Serengeti ungulates, and freshwater sticklebacks) enables new analyses that were computationally impossible at the time. LDP projects help environmental organizations archive data over broad spatial and temporal scales, critical for metrics such as the Watershed Reports or Living Planet Index. LDP also assists non-professionals to archive data; e.g., helping communities archive water quality data on DataStream enables easy comparison to national standards. Data rescue can knit together disparate datasets in relational databases, such as LDP compilation of fish, invertebrate, phytoplankton, and limnology data from the Canadian government’s Turkey Lakes Watershed research. The benefits of data rescue will pay incalculable dividends for decades to come. However, to ensure successful data rescue, we recommend the following: Rules for success: Prevent scope creep. When non-essential data is included in the initial scope of the project or the scope broadens during the rescue process, concluding the project is challenging. It is better to successfully archive a subset of the data than fail to archive any. Involve data collectors and custodians throughout. They provide important contextual knowledge, help decipher obscure data codes and can suggest constraints to values, facilitating validation. Formalizing their knowledge in the metadata ensures future usability. Continually champion projects until the dataset DOI is minted. It is easy for projects to stall in the final stages. Delaying archiving to add more years or variables or to publish one more paper increases the risk of no data being archived. Encourage the use of versioning or embargo periods, available for many repositories, to deposit promptly the core data while enabling later additions and public release. Manage adaptively. Problems like contradictory data, inconsistent use of terms, and heterogeneous collection methods may only be apparent after working with the data. At the LDP, we use high-level checks of progress throughout each rescue project to adaptively manage the scope and personnel needed, informed by our experience in a wide variety of projects. Prevent scope creep. When non-essential data is included in the initial scope of the project or the scope broadens during the rescue process, concluding the project is challenging. It is better to successfully archive a subset of the data than fail to archive any. Involve data collectors and custodians throughout. They provide important contextual knowledge, help decipher obscure data codes and can suggest constraints to values, facilitating validation. Formalizing their knowledge in the metadata ensures future usability. Continually champion projects until the dataset DOI is minted. It is easy for projects to stall in the final stages. Delaying archiving to add more years or variables or to publish one more paper increases the risk of no data being archived. Encourage the use of versioning or embargo periods, available for many repositories, to deposit promptly the core data while enabling later additions and public release. Manage adaptively. Problems like contradictory data, inconsistent use of terms, and heterogeneous collection methods may only be apparent after working with the data. At the LDP, we use high-level checks of progress throughout each rescue project to adaptively manage the scope and personnel needed, informed by our experience in a wide variety of projects. Summary. Environmental data rescue has many benefits, including establishing baselines and change in the environment, enabling new analyses and composite indices, connecting disparate data, facilitating community science, and training the next generation of researchers in open science. However, data rescue projects requires continued and careful management to ensure success, which we defined as the deposition of data in open, accessible and permanent repositories.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.016
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScience and technology studies, Scholarly communication, Open science
Consensus categoriesScholarly communication, Open science
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.772
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0160.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.002
Science and technology studies0.0030.001
Scholarly communication0.0060.079
Open science0.0140.018
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.067
GPT teacher head0.374
Teacher spread0.306 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes3
Has abstractyes

Explore more

Same venueBiodiversity Information Science and StandardsSame topicResearch Data Management PracticesFrench-language works237,207