A novel sea ice floe fragmentation index using Sentinel-2 and AMSR2 satellite data based on machine learning
Bibliographic record
Abstract
• Floe fragmentation index (FFI) quantifies the degree of sea ice floe fragmentation. • FFI revealed floe structural variation even at the same sea ice concentration (SIC). • FFI enables daily floe fragmentation monitoring via microwave and machine learning. • FFI proacted clearly to sea ice loss, even when SIC remained stable. • Rapid increase in FFI was consistently followed by a sharp decline in SIC. Sea ice indices such as sea ice concentration (SIC) play a key role in monitoring climate change. However, it does not fully capture the vulnerability of sea ice to melting, especially under conditions of floe fragmentation. To address this limitation, we introduce a novel metric—the floe fragmentation index (FFI)—designed to quantify the degree of fragmentation of sea ice. We constructed the FFI reference map using high-resolution Sentinel-2 imagery based on k-means clustering and manual editing. This reference was then paired with AMSR2 passive microwave data to train three machine learning models—gradient boosting (GB), random forest (RF), and support vector regression (SVR)—enabling consistent, daily mapping of FFI across the Arctic. FFI increases as sea ice becomes more fragmented. Even under identical SIC conditions (e.g., 50%), the reference FFI captured distinct differences in floe structure, demonstrating its ability to represent fragmentation more explicitly than SIC. In comparison with the reference FFI derived from Sentinel-2, the gradient boosting model demonstrated the best performance, with an R 2 exceeding 0.94 and a root mean square error (RMSE) below 0.22. Since RMSE was computed against the reference FFI, it is expressed in the same unit as FFI, which is dimensionless. To examine differences between SIC and FFI under relatively stable sea ice conditions, we focused on the Laptev Sea—a region where import and export of sea ice are minimal in summer. At the point of sea ice disappearance within this area, no early signs were observed in SIC, whereas FFI did reveal such signals. In particular, under conditions where the sea ice was highly fragmented and thus more likely to drift away or melt, SIC remained close to 100%, while FFI captured the relatively severe fragmentation of the sea ice. The time-series comparison between the FFI and SIC revealed that a rapid increase in FFI, highly exceeding its seasonal tendency, was followed by or occurred simultaneously with a rapid decrease in SIC, which also exceeded twice the magnitude of its seasonal tendency. Ultimately, the FFI complemented SIC by capturing fragmentation changes, potentially allowing earlier detection of SIC change in the melting seasons.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".