A method for the objective selection of landscape‐scale study regions and sites at the national level
Bibliographic record
Abstract
Summary Ecological processes operating on large spatio‐temporal scales are difficult to disentangle with traditional empirical approaches. Alternatively, researchers can take advantage of ‘natural’ experiments, where experimental control is exercised by careful site selection. Recent advances in developing protocols for designing these ‘pseudo‐experiments’ commonly do not consider the selection of the focal region and predictor variables are usually restricted to two. Here, we advance this type of site selection protocol to study the impact of multiple landscape scale factors on pollinator abundance and diversity across multiple regions. Using datasets of geographic and ecological variables with national coverage, we applied a novel hierarchical computation approach to select study sites that contrast as much as possible in four key variables, while attempting to maintain regional comparability and national representativeness. There were three main steps to the protocol: (i) selection of six 100 × 100 km2 regions that collectively provided land cover representative of the national land average, (ii) mapping of potential sites into a multivariate space with axes representing four key factors potentially influencing insect pollinator abundance, and (iii) applying a selection algorithm which maximized differences between the four key variables, while controlling for a set of external constraints. Validation data for the site selection metrics were recorded alongside the collection of data on pollinator populations during two field campaigns. While the accuracy of the metric estimates varied, the site selection succeeded in objectively identifying field sites that differed significantly in values for each of the four key variables. Between‐variable correlations were also reduced or eliminated, thus facilitating analysis of their separate effects. This study has shown that national datasets can be used to select randomized and replicated field sites objectively within multiple regions and along multiple interacting gradients. Similar protocols could be used for studying a range of alternative research questions related to land use or other spatially explicit environmental variables, and to identify networks of field sites for other countries, regions, drivers and response taxa in a wide range of scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.051 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.013 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".