Small but visible: Predicting rare bryophyte distribution and richness patterns using remote sensing-based ensembles of small models
Bibliographic record
Abstract
In Canadian boreal forests, bryophytes represent an essential component of biodiversity and play a significant role in ecosystem functioning. Despite their ecological importance and sensitivity to disturbances, bryophytes are overlooked in conservation strategies due to knowledge gaps on their distribution, which is known as the Wallacean shortfall. Rare species deserve priority attention in conservation as they are at a high risk of extinction. This study aims to elaborate predictive models of rare bryophyte species in Canadian boreal forests using remote sensing-derived predictors in an Ensemble of Small Models (ESMs) framework. We hypothesize that high ESMs-based prediction accuracy can be achieved for rare bryophyte species despite their low number of occurrences. We also assess if there is a spatial correspondence between rare and overall bryophyte richness patterns. The study area is located in western Quebec and covers 72,292 km2. We selected 52 bryophyte species with <30 occurrences from a presence-only database (214 species, 389 plots in total). ESMs were built from Random Forest and Maxent techniques using remote sensing-derived predictors related to topography and vegetation. Lee's L statistic was used to assess and map the spatial relationship between rare and overall bryophyte richness patterns. ESMs yielded poor to excellent prediction accuracy (AUC > 0.5) for 73% of the modeled species, with AUC values > 0.8 for 19 species, which confirmed our hypothesis. In fact, ESMs provided better predictions for the rarest bryophytes. Likewise, our study revealed a spatial concordance between rare and overall bryophyte richness patterns in different regions of the study area, which have important implications for conservation planning. This study demonstrates the potential of remote sensing for assessing and making predictions on inconspicuous and rare species across the landscape and lays the basis for the eventual inclusion of bryophytes into sustainable development planning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".