Leveraging Open‐Source Geographic Databases to Enhance the Representation of Landscape Heterogeneity in Ecological Models
Bibliographic record
Abstract
Wildlife abundance and movement are strongly impacted by landscape heterogeneity, especially in cities which are among the world's most heterogeneous landscapes. Nonetheless, current global land cover maps, which are used as a basis for large-scale spatial ecological modeling, represent urban areas as a single, homogeneous, class. This often requires urban ecologists to rely on geographic resources from local governments, which are not comparable between cities and are not available in underserved countries, limiting the spatial scale at which urban conservation issues can be tackled. The recent expansion of community-based geographic databases, for example, OpenStreetMap (OSM), represents an opportunity for ecologists to generate large-scale maps geared toward their specific research needs. However, computational differences in language and format, and the high diversity of information within, limit the access to these data. We provide a framework, using R, to extract geographic features from the OSM database, classify, and integrate them into global land cover maps. The framework includes an exhaustive list of OSM features describing urban and peri-urban landscapes and is validated by quantifying the completeness of the OSM features characterized, and the accuracy of its final output in 34 cities in North America. We portray its application as the basis for generating landscape variables for ecological analysis by using the OSM-enhanced map to generate an urbanization index, and subsequently analyze the spatial occupancy of six mammals throughout Chicago, Illinois, USA. The OSM features characterized had high completeness values for impervious land cover classes (50%-100%). The final output, the OSM-enhance map, provided an 89% accurate representation of the landscape at 30m resolution. The OSM-derived urbanization index outperformed other global spatial data layers in the spatial occupancy analysis and concurred with previously seen local response trends, whereby lagomorphs and squirrels responded positively to urbanization, while skunks, raccoons, opossums, and deer responded negatively. This study provides a roadmap for ecologists to leverage the fine resolution of open-source geographic databases and apply it to spatial modeling by generating research-specific landscape variables. As our occupancy results show, using context-specific maps can improve modeling outputs and reduce uncertainty, especially when trying to understand anthropogenic impacts on wildlife populations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.035 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".