Comparing wildlife habitat suitability models based on expert opinion with camera trap detections
Bibliographic record
Abstract
Expert knowledge is used in the development of wildlife habitat suitability models (HSMs) for management and conservation decisions. However, the consistency of such models has been questioned. Focusing on 1 method for elicitation, the analytic hierarchy process, we generated expert-based HSMs for 4 felid species: 2 forest specialists (ocelot [Leopardus pardalis] and margay [Leopardus wiedii]) and 2 habitat generalist species (Pampas cat [Leopardus colocola] and puma [Puma concolor]). Using these HSMs, species detections from camera-trap surveys, and generalized linear models, we assessed the effect of study species and expert attributes on the correspondence between expert models and camera-trap detections. We also examined whether aggregation of participant responses and iterative feedback improved model performance. We ran 160 HSMs and found that models for specialist species showed higher correspondence with camera-trap detections (AUC [area under the receiver operating characteristic curve] >0.7) than those for generalists (AUC < 0.7). Model correspondence increased as participant years of experience in the study area increased, but only for the understudied generalist species, Pampas cat (β = 0.024 [SE 0.007]). No other participant attribute was associated with model correspondence. Feedback and revision of models improved model correspondence, and aggregating judgments across multiple participants improved correspondence only for specialist species. The average correspondence of aggregated judgments increased as group size increased but leveled off after 5 experts for all species. Our results suggest that correspondence between expert models and empirical surveys increases as habitat specialization increases. We encourage inclusion of participants knowledgeable of the study area and model validation for expert-based modeling of understudied and generalist species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".