The forests of the midwestern United States at Euro-American settlement: Spatial and physical structure based on contemporaneous survey data
Bibliographic record
Abstract
We present gridded 8 km-resolution data products of the estimated stem density, basal area, and biomass of tree taxa at Euro-American settlement of the midwestern United States during the middle to late 19th century for the states of Minnesota, Wisconsin, Michigan, Illinois, and Indiana. The data come from settlement-era Public Land Survey (PLS) data (ca. 0.8-km resolution) of trees recorded by land surveyors. The surveyor notes have been transcribed, cleaned, and processed to estimate stem density, basal area, and biomass at individual points. The point-level data are aggregated within 8 km grid cells and smoothed using a generalized additive statistical model that accounts for zero-inflated continuous data and provides approximate Bayesian uncertainty estimates. The statistical modeling smooths out sharp spatial features (likely arising from statistical noise) within areas smaller than about 200 km2. Based on this modeling, presettlement Midwestern landscapes supported multiple dominant species, vegetation types, forest types, and ecological formations. The prairies, oak savannas, and forests each had distinctive structures and spatial distributions across the domain. Forest structure varied from savanna (averaging 27 Mg/ha biomass) to northern hardwood (104 Mg/ha) and mesic southern forests (211 Mg/ha). The presettlement forests were neither unbroken and massively-statured nor dominated by young forests constantly structured by broad-scale disturbances such as fire, drought, insect outbreaks, or hurricanes. Most forests were structurally between modern second growth and old growth. We expect the data product to be useful as a baseline for investigating how forest ecosystems have changed in response to the last several centuries of climate change and intensive Euro-American land use and as a calibration dataset for paleoecological proxy-based reconstructions of forest composition and structure for earlier time periods. The data products (including raw and smoothed estimates at the 8-km scale) are available at the LTER Network Data Portal as version 1.0.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".