MétaCan
Menu
Back to cohort
Record W4408429610 · doi:10.5194/egusphere-egu25-15864

The Sentinels EOPF Toolkit: Driving Community Adoption of the Zarr data format for Copernicus Sentinel Data

2025· preprint· en· W4408429610 on OpenAlexaff
Sabrina H. Szeto, Julia Wagemann, Emmanuel Mathot, James Banting

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldComputer Science
TopicAdvanced Computational Techniques and Applications
Canadian institutionsPositive Living North
Fundersnot available
KeywordsCopernicusComputer scienceCloud computingWorkflowWorld Wide WebGeospatial analysisData scienceData accessKey (lock)DatabaseRemote sensingComputer securityGeography

Abstract

fetched live from OpenAlex

The Standard Archive Format for Europe (SAFE) specification has been the established approach to publishing Copernicus Sentinel data products for over a decade. While SAFE has pushed the ecosystem forward through new ways to search and access the data, it is not ideal for processing large volumes of data using cloud computing. Over the last few years, data standards like STAC and cloud-native data formats like Zarr and COGs have revolutionised how scientific communities work with large-scale geospatial data and are becoming a key component of new data spaces, especially for cloud-based systems.The ESA Copernicus Earth Observation Processor Framework (EOPF) will be providing access to “live” sample data from the Copernicus Sentinel missions -1, -2 and -3 in the new Zarr data format. This set of reprocessed data allows users to try out accessing and processing data in the new format and experiencing the benefits thereof with their own workflows.This presentation introduces a community-driven toolkit that facilitates the adoption of the Zarr data format for Copernicus Sentinel data. The creation of this toolkit was driven by several motivating questions: What common challenges do users face and how can we help them overcome them? What resources would make it easier for Sentinel data users to use the new Zarr data format? How can we foster a community of users who will actively contribute to the creation of this toolkit and support each other? The Sentinels EOPF Toolkit team, comprising Development Seed, SparkGeo and thriveGEO, together with a group of champion users (early-adopters), are creating a set of Jupyter Notebooks and plug-ins that showcase the use of Zarr format Sentinel data for applications across multiple domains. In addition, community engagement activities such as a notebook competition and social media outreach will bring Sentinel users together and spark interaction with the new data format in a creative yet supportive environment. Such community and user adoption efforts are necessary in order to overcome adoption and uptake barriers and to build up trust and excitement to try out new technologies and new developments around data spaces.In addition to introducing the Sentinels EOPF Toolkit, this presentation will also highlight lessons learned from working closely with users on barriers they face in adopting the new Zarr format and how to address them.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.042
metaresearch head score (Gemma)0.064
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.042
Threshold uncertainty score0.223

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0420.064
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0040.003
Science and technology studies0.0030.003
Scholarly communication0.0070.013
Open science0.0050.023
Research integrity0.0020.005
Insufficient payload (model declined to judge)0.0110.013

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.127
GPT teacher head0.381
Teacher spread0.254 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same topicAdvanced Computational Techniques and ApplicationsFrench-language works237,207