Accuracy assessment of global land cover products derived from satellite data
Bibliographic record
Abstract
Information on land cover distribution at regional and global scales has become fundamental for studying global changes affecting ecological and climatic systems. The remote sensing community has responded to this increased interest by improving data quality and methodologies for extracting land cover information. However, in addition to the advantages provided by satellite products, certain limitations exist that need to be objectively quantified and clearly communicated to users so that they can make informed decisions on if and how land cover products should be used. Accuracy assessment is the procedure used to quantify product quality. Some aspects of accuracy assessment for evaluating four global land cover maps over Canada are discussed in this paper. Attempts are made to quantify limiting factors resulting from the coarse spatial resolution of data used for generating land cover information at regional and global levels. Sub-pixel fractional error matrices are introduced as a more appropriate way for assessing the accuracy of mixed pixels. For classification with coarse spatial resolution data, limitations of the classification method produce a maximum achievable accuracy defined as the average dominant land cover fraction of the mapped area. Relationships among spatial resolution, landscape heterogeneity and thematic resolution were studied and reported. Other factors that can affect accuracy, such as misregistration and legend conversion, are also discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".