A Comparison of Deep Transfer Learning Methods for Land Use and Land Cover Classification
Bibliographic record
Abstract
The pace of Land Use/Land Cover (LULC) change has accelerated due to population growth, industrialization, and economic development. To understand and analyze this transformation, it is essential to examine changes in LULC meticulously. LULC classification is a fundamental and complex task that plays a significant role in farming decision making and urban planning for long-term development in the earth observation system. Recent advances in deep learning, transfer learning, and remote sensing technology have simplified the LULC classification problem. Deep transfer learning is particularly useful for addressing the issue of insufficient training data because it reduces the need for equally distributed data. In this study, thirty-nine deep transfer learning models were systematically evaluated alongside multiple deep transfer learning models for LULC classification using a consistent set of criteria. Our experiments will be conducted under controlled conditions to provide valuable insights for future research on LULC classification using deep transfer learning models. Among our models, ResNet50, EfficientNetV2B0, and ResNet152 were the top performers in terms of kappa and accuracy scores. ResNet152 required three times longer training time than EfficientNetV2B0 on our test computer, while ResNet50 took roughly twice as long. ResNet50 achieved an overall f1-score of 0.967 on the test set, with the Highway class having the lowest score and the Sea Lake class having the highest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.006 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".