Köppen meets Neural Network: Revision of the Köppen Climate Classification by Neural Networks
Bibliographic record
Abstract
Climate change and development of data-oriented methods are appealing for new climate classification schemes. Based on the most widely used Köppen-Geiger scheme, this article proposes a neural network based climate classification method from a data science perspective. In conventional schemes, empirically handcrafted rules are used to divide climate data into climate types, resulting in certain defects. In the proposed method, a machine learning mechanism is employed to do the task. Specifically, the method first trains a convolutional neural network to fit climate data to land cover conditions, then extracts features from the trained network and finally uses a self-organizing map to cluster land pixels on the extracted features. The method is applied to cluster global land represented by 66,501 pixels (each covers 0.5 latitude degree × 0.5 longitude degree) using 2020 land cover data and 1991-2020 climate normals, and a 4 × 3 × 2 hexagonal self-organizing map clusters the land pixels into twenty-four climate types. By Kappa statistics, the obtained scheme shows good agreement with the Köppen-Geiger and Köppen-Trewartha schemes. In addition, our scheme addresses some issues of the Köppen schemes, suggests new climate types such as As (severe dry-wet season) and Fw (arctic desert), and identifies the highland group H without input of elevation. The proposed method is expected as an intelligent tool to monitor changes in the global climate pattern and to discover new climate types of interest that possibly emerge in the future. It may also be valuable for bio-ecology communities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.006 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".