From need to product: A methodology for completing a land cover map of Canada from Landsat
Bibliographic record
Abstract
Abstract. Despite its very large territory and the best Landsat archive in the world, Canada has made very limited use of Landsat data for land cover mapping. The primary difficulty has been the prohibitive cost of information extraction and the earlier (and now overcome for Landsat-7 enhanced thematic mapper plus data) high cost of data purchase. The solution to this remaining obstacle lies in decreasing the cost of Landsat data processing and analysis while ensuring the high quality of the extracted information. In this paper, we present an efficient and effective approach to mapping land cover in Canada from Landsat thematic mapper data (single or multiple satellites). The key features of this approach are an increase in the ratio of computer to human analysis and automation for high data volume or large area processing. However, it is essential that the final product quality not suffer because of the greater reliance on computer processing, thus the algorithm performance becomes critical. We describe the overall approach, discuss key challenges, explain the principles behind key algorithms developed to respond to the challenges, present evidence demonstrating the effectiveness of these algorithms in a boreal landscape setting, and consider implementation issues. With a processing system developed to handle large numbers (tens to hundreds) of Landsat scenes, which incorporates most of the algorithms discussed here, the stage is nearly set for large-scale processing leading to a Landsat-based land cover classification product(s) for Canada. Résumé. En dépit de l’étendue de son territoire et de la disponibilité de la meilleure archive Landsat au monde, le Canada a jusqu’à maintenant très peu fait usage des données Landsat pour les besoins de la cartographie du couvert. La difficulté
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".