Assessing the Temporal Stability of the Accuracy of a Time Series of Burned Area Products
Bibliographic record
Abstract
Temporal stability, defined as the change of accuracy through time, is one of the validation aspects required by the Committee on Earth Observation Satellites’ Land Product Validation Subgroup. Temporal stability was evaluated for three burned area products: MCD64, Globcarbon, and fire_cci. Traditional accuracy measures, such as overall accuracy and omission and commission error ratios, were computed from reference data for seven years (2001–2007) in seven study sites, located in Angola, Australia, Brazil, Canada, Colombia, Portugal, and South Africa. These accuracy measures served as the basis for the evaluation of temporal stability of each product. Nonparametric tests were constructed to assess different departures from temporal stability, specifically a monotonic trend in accuracy over time (Wilcoxon test for trend), and differences in median accuracy among years (Friedman test). When applied to the three burned area products, these tests did not detect a statistically significant temporal trend or significant differences among years, thus, based on the small sample size of seven sites, there was insufficient evidence to claim these products had temporal instability. Pairwise Wilcoxon tests comparing yearly accuracies provided a measure of the proportion of year-pairs with significant differences and these proportions of significant pairwise differences were in turn used to compare temporal stability between BA products. The proportion of year-pairs with different accuracy (at the 0.05 significance level) ranged from 0% (MCD64) to 14% (fire_cci), computed from the 21 year-pairs available. In addition to the analysis of the three real burned area products, the analyses were applied to the accuracy measures computed for four hypothetical burned area products to illustrate the properties of the temporal stability analysis for different hypothetical scenarios of change in accuracy over time. The nonparametric tests were generally successful at detecting the different types of temporal instability designed into the hypothetical scenarios. The current work presents for the first time methods to quantify the temporal stability of BA product accuracies and to alert product end-users that statistically significant temporal instabilities exist. These methods represent diagnostic tools that allow product users to recognize the potential confounding effect of temporal instability on analysis of fire trends and allow map producers to identify anomalies in accuracy over time that may lead to insights for improving fire products. Additionally, we suggest temporal instabilities that could hypothetically appear, caused by for example by failures or changes in sensor data or classification algorithms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".