Inter-comparison and evaluation of Arctic sea ice type products
Bibliographic record
Abstract
Abstract. Arctic sea ice type (SITY) variation is a sensitive indicator of climate change. However, systematic inter-comparison and analysis for SITY products are lacking. This study analysed eight daily SITY products from five retrieval approaches covering the winters of 1999–2019, including purely radiometer-based (C3S-SITY), scatterometer-based (KNMI-SITY and IFREMER-SITY) and combined ones (OSISAF-SITY and Zhang-SITY). These SITY products were inter-compared against a weekly sea ice age product (i.e. NSIDC-SIA – National Snow and Ice Data Center sea ice age) and evaluated with five synthetic aperture radar (SAR) images. The average Arctic multiyear ice (MYI) extent difference between the SITY products and NSIDC-SIA varies from -1.32×106 to 0.49×106 km2. Among them, KNMI-SITY and Zhang-SITY in the QuikSCAT (QSCAT) period (2002–2009) agree best with NSIDC-SIA and perform the best, with the smallest bias of -0.001×106 km2 in first-year ice (FYI) extent and -0.02×106 km2 in MYI extent. In the Advanced Scatterometer (ASCAT) period (2007–2019), KNMI-SITY tends to overestimate MYI (especially in early winter), whereas Zhang-SITY and IFREMER-SITY tend to underestimate MYI. C3S-SITY performs well in some early winter cases but exhibits large temporal variabilities like OSISAF-SITY. Factors that could impact performances of the SITY products are analysed and summarized. (1) The Ku-band scatterometer generally performs better than the C-band scatterometer for SITY discrimination, while the latter sometimes identifies FYI more accurately, especially when surface scattering dominates the backscatter signature. (2) A simple combination of scatterometer and radiometer data is not always beneficial without further rules of priority. (3) The representativeness of training data and efficiency of classification are crucial for SITY classification. Spatial and temporal variation in characteristic training datasets should be well accounted for in the SITY method. (4) Post-processing corrections play important roles and should be considered with caution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".