Urban Land Use and Land Cover Change Analysis Using Random Forest Classification of Landsat Time Series
Bibliographic record
Abstract
Efficient implementation of remote sensing image classification can facilitate the extraction of spatiotemporal information for land use and land cover (LULC) classification. Mapping LULC change can pave the way to investigate the impacts of different socioeconomic and environmental factors on the Earth’s surface. This study presents an algorithm that uses Landsat time-series data to analyze LULC change. We applied the Random Forest (RF) classifier, a robust classification method, in the Google Earth Engine (GEE) using imagery from Landsat 5, 7, and 8 as inputs for the 1985 to 2019 period. We also explored the performance of the pan-sharpening algorithm on Landsat bands besides the impact of different image compositions to produce a high-quality LULC map. We used a statistical pan-sharpening algorithm to increase multispectral Landsat bands’ (Landsat 7–9) spatial resolution from 30 m to 15 m. In addition, we checked the impact of different image compositions based on several spectral indices and other auxiliary data such as digital elevation model (DEM) and land surface temperature (LST) on final classification accuracy based on several spectral indices and other auxiliary data on final classification accuracy. We compared the classification result of our proposed method and the Copernicus Global Land Cover Layers (CGLCL) map to verify the algorithm. The results show that: (1) Using pan-sharpened top-of-atmosphere (TOA) Landsat products can produce more accurate results for classification instead of using surface reflectance (SR) alone; (2) LST and DEM are essential features in classification, and using them can increase final accuracy; (3) the proposed algorithm produced higher accuracy (94.438% overall accuracy (OA), 0.93 for Kappa, and 0.93 for F1-score) than CGLCL map (84.4% OA, 0.79 for Kappa, and 0.50 for F1-score) in 2019; (4) the total agreement between the classification results and the test data exceeds 90% (93.37–97.6%), 0.9 (0.91–0.96), and 0.85 (0.86–0.95) for OA, Kappa values, and F1-score, respectively, which is acceptable in both overall and Kappa accuracy. Moreover, we provide a code repository that allows classifying Landsat 4, 5, 7, and 8 within GEE. This method can be quickly and easily applied to other regions of interest for LULC mapping.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".