An integrated framework for assessing the accuracy of GEOBIA land cover products
Bibliographic record
Abstract
Vector-based landcover (LC) maps derived from GEographic Object-Based Image Analysis (GEOBIA) are increasingly replacing the traditional raster maps from per-pixel classification, but our strategies for assessing their quality are not yet fully developed. We contend that a complete accuracy assessment of a vector LC map must provide answers to the following questions: (1) What is the proportion of area assigned to each LC class that is actually covered by that class? (2) How does the area wrongly assigned to a class get distributed into the other classes? If we were flying at a low altitude over any given polygon, what is the likelihood that we would agree that (3) the LC class best representing the interior of the polygon is the one appearing on the map; (4) the area enclosed by the polygon can be seen as a self-contained unit or patch; (5) there are no regions, either next to the outside of the polygon or on its inside, that would have better be included in the polygon or excluded from it; and (6) the outline of the polygon (excluding parts affected by 5) follows reasonably well the LC transitions we appreciate from air? Questions 1 and 2 can be answered using a confusion matrix, but not the rest. We discuss the conceptual foundations of our integrated object-based approach to accuracy assessment, and demonstrate its implementation for a wall to wall vector LC map of Alberta, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.005 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".