Special Observing Period (SOP) Data for the Year of Polar Prediction site Model Intercomparison Project (YOPPsiteMIP)
Bibliographic record
Abstract
Abstract. The rapid changes occurring in the polar regions require an improved understanding of the processes that are driving the changes. At the same time increased human activities, such as marine navigation, resource exploitation, aviation, commercial fishing, and tourism, require reliable and relevant information. One of the primary goals of the World Meteorological Organization’s Year of Polar Prediction (YOPP) Project is to improve the accuracy of numerical weather prediction (NWP) at high latitudes. During YOPP, two Canadian observatories were commissioned and equipped with new ground-based instruments for enhanced meteorological and system process observations that are considered to be “supersites” for addressing YOPP objectives, while other pre-existing supersites in Canada, the United States, Norway, Finland and Russia provided data from ongoing long-term observing programs. Data from these seven supersites were amalgamated and are being used to evaluate NWP systems from several international forecast centers and to perform meteorological process studies with the aim of improving NWP performance in the Polar Regions. In order to increase data useability and station interoperability, novel Merged Observatory Data Files (MODFs) have been created for these seven international supersites over two Special Observing Periods (February to March 2018 and July to September 2018). All observations collected at the seven supersites were compiled into this new standardized NetCDF MODF format, simplifying the process of conducting pan-Arctic NWP verification and process evaluation studies. This paper describes the seven Arctic YOPP supersites, data collection and processing methods, and the novel MODF format and output files, which together comprise the observational contribution to the associated model intercomparison effort, termed YOPP supersite Model Intercomparison Project (YOPPsiteMIP). All YOPPsiteMIP MODFs are publicly accessible via the YOPP Data Portal (Whitehorse: https://doi.org/10.21343/a33e-j150, Iqaluit: https://doi.org/10.21343/yrnf-ck57, Sodankylä: https://doi.org/10.21343/m16p-pq17, Utqiaġvik: https://doi.org/10.21343/a2dx-nq55, Tiksi: https://doi.org/10.21343/5bwn-w881, Ny-Ålesund: https://doi.org/10.21343/y89m-6393, Eureka: https://doi.org/10.21343/r85j-tc61), hosted by MET Norway, with corresponding output from NWP models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".