Comparison of NASA Team2 and AES-york ice concentration algorithms against operational ice charts from the Canadian ice service
Bibliographic record
Abstract
Ice concentration retrieved from spaceborne passive-microwave observations is a prime input to operational sea-ice-monitoring programs, numerical weather prediction models, and global climate models. Atmospheric Environment Service (AES)-York and the Enhanced National Aeronautics and Space Administration Team (NT2) are two algorithms that calculate ice concentration from SpecialSensor Microwave/Imager observations. This paper furnishes a comparison between ice concentrations (total, thin, and thick types) output from NT2 and AES-York algorithms against the corresponding estimates from the operational analysis of Radarsat images in the Canadian Ice Service (CIS). A new data fusion technique, which incorporates the actual sensor's footprint, was developed to facilitate this study. Results have shown that the NT2 and AES-York algorithms underestimate total ice concentration by 18.35% and 9.66% concentration counts on average, with 16.8% and 15.35% standard deviation, respectively. However, the retrieved concentrations of thin and thick ice are in much more discrepancy with the operational CIS estimates when either one of these two types dominates the viewing area. This is more likely to occur when the total ice concentration approaches 100%. If thin and thick ice types coexist in comparable concentrations, the algorithms' estimates agree with CIS's estimates. In terms of ice concentration retrieval, thin ice is more problematic than thick ice. The concept of using a single tie point to represent a thin ice surface is not realistic and provides the largest error source for retrieval accuracy. While AES-York provides total ice concentration in slightly more agreement with CIS's estimates, NT2 provides better agreement in retrieving thin and thick ice concentrations
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".