Geospatial Data Preservation Prime
Bibliographic record
Abstract
This primer is one in a series of Operational Policy documents being developed by GeoConnections. It is intended to inform Canadian Geospatial Data Infrastructure (CGDI) stakeholders about the nature and scope of digital geospatial data archiving and preservation and the realities, challenges and good practices of related operational policies. Burgeoning growth of online geospatial applications and the deluge of data, combined with the growing complexity of archiving and preserving digital data, has revealed a significant gap in the operational policy coverage for the Canadian geospatial data infrastructure (CGDI). Currently there is no commonly accepted guidance for CGDI stakeholders wishing or mandated to preserve their geospatial data assets for long-term access and use. More specifically, there is little or no guidance available to inform operational policy decisions on how to manage, preserve and provide access to a digital geospatial data collection. The preservation of geospatial data over a period of time is especially important when datasets are required to inform modeling applications such as climate change impact predictions, flood forecasts and land use management. Furthermore, data custodians may have both a legal and moral responsibility to implement effective archiving and preservation programs. Based on research and analysis of the Canadian legislative framework and current international practices in digital data archiving and preservation, this primer provides guidance on the factors to be considered and the steps to be taken in planning and implementing a data archiving and preservation program. It describes an approach to establishing a geospatial data archives based on good practices from the literature and Canadian case studies. This primer will provide CGDI stakeholders with information on how to incorporate archiving and preservation considerations into an effective data management process that covers the entire life cycle (DCC, 2013) (LAC, 2006) of their geospatial data assets (i.e., creation and receipt, distribution, use, maintenance, and disposition. It is intended to inform CGDI stakeholders on the importance of long term data preservation, and provide them with the information and tools required to make policy decisions for creating an archives and preserving digital geospatial data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.005 | 0.003 |
| Scholarly communication | 0.008 | 0.007 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.073 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".