Rescuing the Past to Prepare for the Future: Environmental Data Rescue as a Key Activity of Open Science
Notice bibliographique
Résumé
The need for data rescue. Data are dying all around us: when they are stored in inaccessible or fragile repositories, in proprietary formats, on defunct media or without complete metadata. In environmental science, such data loss has high societal and scientific costs. For example, failure to archive ecological data focused on the Exxon Valdez oil spill represents an estimated loss of >100 million USD (Bledsoe et al. 2022). Loss of environmental data can also represent erasure of historic baselines that can never be recollected. The Living Data Project ( LDP ). This award-winning program helps scientists and organizations archive valuable “legacy” datasets, making them permanently open and accessible (Bledsoe et al. 2022). We pair data custodians with graduate students who we train in reproducible data management and provide mentoring by postdoctoral researchers and faculty and assistance by undergraduates. By training Canada’s next generation of scientists, we aim to ensure that data are never lost again. Over the last five years, we have trained 305 students, across 93 internships, rescuing > 2000 years of data. LDP student working groups then analyze archived data to investigate pressing biodiversity questions and develop open access teaching resources. Benefits of data rescue. LDP-rescued datasets, many stretching back into the last century, provide missing historical baselines that enable the detection of environmental change – such as change in lake ice formation and melt over the last 150 years. Rescue of historical datasets, like surveys of urban birds or plant communities collected five decades ago, sets the stage for contemporary re-surveys. Combining rescued data with contemporary data can also yield discoveries such as legacy effects of pesticides on stream health (Sugden et al. 2025). LDP rescue of classic studies in ecological and evolutionary theory (Darwin’s Finches, Serengeti ungulates, and freshwater sticklebacks) enables new analyses that were computationally impossible at the time. LDP projects help environmental organizations archive data over broad spatial and temporal scales, critical for metrics such as the Watershed Reports or Living Planet Index. LDP also assists non-professionals to archive data; e.g., helping communities archive water quality data on DataStream enables easy comparison to national standards. Data rescue can knit together disparate datasets in relational databases, such as LDP compilation of fish, invertebrate, phytoplankton, and limnology data from the Canadian government’s Turkey Lakes Watershed research. The benefits of data rescue will pay incalculable dividends for decades to come. However, to ensure successful data rescue, we recommend the following: Rules for success: Prevent scope creep. When non-essential data is included in the initial scope of the project or the scope broadens during the rescue process, concluding the project is challenging. It is better to successfully archive a subset of the data than fail to archive any. Involve data collectors and custodians throughout. They provide important contextual knowledge, help decipher obscure data codes and can suggest constraints to values, facilitating validation. Formalizing their knowledge in the metadata ensures future usability. Continually champion projects until the dataset DOI is minted. It is easy for projects to stall in the final stages. Delaying archiving to add more years or variables or to publish one more paper increases the risk of no data being archived. Encourage the use of versioning or embargo periods, available for many repositories, to deposit promptly the core data while enabling later additions and public release. Manage adaptively. Problems like contradictory data, inconsistent use of terms, and heterogeneous collection methods may only be apparent after working with the data. At the LDP, we use high-level checks of progress throughout each rescue project to adaptively manage the scope and personnel needed, informed by our experience in a wide variety of projects. Prevent scope creep. When non-essential data is included in the initial scope of the project or the scope broadens during the rescue process, concluding the project is challenging. It is better to successfully archive a subset of the data than fail to archive any. Involve data collectors and custodians throughout. They provide important contextual knowledge, help decipher obscure data codes and can suggest constraints to values, facilitating validation. Formalizing their knowledge in the metadata ensures future usability. Continually champion projects until the dataset DOI is minted. It is easy for projects to stall in the final stages. Delaying archiving to add more years or variables or to publish one more paper increases the risk of no data being archived. Encourage the use of versioning or embargo periods, available for many repositories, to deposit promptly the core data while enabling later additions and public release. Manage adaptively. Problems like contradictory data, inconsistent use of terms, and heterogeneous collection methods may only be apparent after working with the data. At the LDP, we use high-level checks of progress throughout each rescue project to adaptively manage the scope and personnel needed, informed by our experience in a wide variety of projects. Summary. Environmental data rescue has many benefits, including establishing baselines and change in the environment, enabling new analyses and composite indices, connecting disparate data, facilitating community science, and training the next generation of researchers in open science. However, data rescue projects requires continued and careful management to ensure success, which we defined as the deposition of data in open, accessible and permanent repositories.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,016 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,002 |
| Études des sciences et des technologies | 0,003 | 0,001 |
| Communication savante | 0,006 | 0,079 |
| Science ouverte | 0,014 | 0,018 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».