Goldilocks and the Raster Grid: Selecting Scale when Evaluating Conservation Programs
Bibliographic record
Abstract
Access to high quality spatial data raises fundamental questions about how to select the appropriate scale and unit of analysis. Studies that evaluate the impact of conservation programs have used multiple scales and areal units: from 5x5 km grids; to 30m pixels; to irregular units based on land uses or political boundaries. These choices affect the estimate of program impact. The bias associated with scale and unit selection is a part of a well-known dilemma called the modifiable areal unit problem (MAUP). We introduce this dilemma to the literature on impact evaluation and then explore the tradeoffs made when choosing different areal units. To illustrate the consequences of the MAUP, we begin by examining the effect of scale selection when evaluating a protected area in Mexico using real data. We then develop a Monte Carlo experiment that simulates a conservation intervention. We find that estimates of treatment effects and variable coefficients are only accurate under restrictive circumstances. Under more realistic conditions, we find biased estimates associated with scale choices that are both too large or too small relative to the data generating process or decision unit. In our context, the MAUP may reflect an errors in variables problem, where imprecise measures of the independent variables will bias the coefficient estimates toward zero. This problem may be pronounced at small scales of analysis. Aggregation may reduce this bias for continuous variables, but aggregation exacerbates bias when using a discrete measure of treatment. While we do not find a solution to these issues, even though treatment effects are generally underestimated. We conclude with suggestions on how researchers might navigate their choice of scale and aerial unit when evaluating conservation policies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".