China's annual forest age dataset at 30 m spatial resolution from 1986 to 2022
Bibliographic record
Abstract
This dataset presents China’s Annual Forest Age (CAFA) at 30-m resolution from 1986 to 2025 (Version 4.0). It was derived by merging forest disturbance detection using Landsat data and age mapping of undisturbed forests using machine learning methods based on forest height, climate, terrain, and Landsat data (Shang et al., 2024; Shang et al., 2023; Lin et al., 2023). The forest extent was determined by the CLCD forest cover dataset (Yang et al. 2021). Areas classified as non-forest are set to -1, and the forest age for the year in which a disturbance occurred is set to 0. We welcome any feedback on our data for future updates!Update HistoryVersion 1.0: First release of China’s 2019 forest age data (Shang et al., 2023).Version 1.1: Updated forest mask with the CLCD forest cover dataset.Version 1.2: Optimized machine learning model inputs to improve age estimation for undisturbed forests, particularly in Northeast and Southwest China.Version 1.3: Enhanced forest disturbance detection using spatial information.Version 1.4: Expanded annual forest age data to 1986–2022 with refined disturbance detection.Version 1.5: Optimized forest disturbance detection through bidirectional monitoring.Version 1.6: Added forest ages before their first forest disturbance.Version 2.0: Second release of China’s annual forest age (CAFA) dataset from 1986 to 2022 (Shang et al., 2025).Version 3.0: Extended annual forest age coverage to 2023–2025.Version 4.0: Updated forest ages by improvinig forest disturbance detection with optimal topographic correction (Yang et al., 2026).NoticePlease click Google Drive to download the full CAFA dataset from 1986 to 2025.Emails: Rong Shang (rongshang90@gmail.com, https://www.researchgate.net/profile/Rong-Shang), Jing M. Chen (jing.chen@utoronto.ca).CitationsShang R., Lin, X. Chen J.M., et al.,(2025), China's annual forest age dataset at a 30 m spatial resolution from 1986 to 2022. Earth System Science Data 17, 3219–3241. [Link]Shang R., Chen J.M., Xu M., et al.,(2023), China's current forest age structure will lead to weakened carbon sinks in the near future. The Innovation 4(6),100515. [Link]Lin, X., Shang, R.*, Chen, J.M., et al.,(2023), High-resolution forest age mapping based on forest height maps derived from GEDI and ICESat-2 space-borne lidar data. Agricultural and Forest Meteorology 339, 109592. [Link]Yang, Z., Shang, R.*, ... , Chen, J.M., (2026), Quantifying the efficacy of topographic correction for forest disturbance monitoring using Landsat time series. Under review.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".