Predicting smallmouth bass (<i>Micropterus dolomieu</i>) occurrence across North America under climate change: a comparison of statistical approaches
Bibliographic record
Abstract
Smallmouth bass (Micropterus dolomieu) is a warm-water fish species that is native to central and eastern North America. Climate change scenarios predict further extension northward of suitable habitat for smallmouth bass, which may negatively affect native fish species. We developed and compared predictive models of the distribution of bass in North America using four statistical approaches: logistic regression, classification tree, discriminant analysis, and artificial neural networks. We collected 4181 geo-referenced records of smallmouth bass occurrence and matched them with climate data. Artificial neural networks performed the best with the highest sensitivity (correctly predicting species presence) and specificity (correctly predicting absence), followed by discriminant analysis. Artificial neural networks indicated that winter air temperatures were the most important predictors of smallmouth bass occurrence, whereas the other approaches indicated that summer air temperatures were the best predictors of bass occurrence. Logistic regression and classification tree exhibited very low sensitivity, but very high specificity as a result of the large proportion of absences within the data set. Business-as-usual climate change scenarios suggest that smallmouth bass are expected to have suitable thermal habitat throughout most of Canada and the continental United States by 2100.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".