Selecting traits that explain species–environment relationships: a generalized linear mixed model approach
Bibliographic record
Abstract
Abstract Question Quantification of the effect of species traits on the assembly of communities is challenging from a statistical point of view. A key question is how species occurrence and abundance can be explained by the trait values of the species and the environmental values at the sites. Methods Using a sites × species abundance table, a site × environment data table and a species × trait data table, we address the above question using a novel generalized linear mixed model (GLMM) approach. TheGLMMovercomes problems of pseudo‐replication and heteroscedastic variance by including sites and species as random factors. The method is equally applicable to presence–absence data as to count and multinomial data. We present a tiered forward selection approach for obtaining a parsimonious model and compare the results with alternative methods (the fourth corner method andRLQordination). Results We illustrate the approach on a presence–absence version on two data sets. In theDuneMeadow data, species presence is parsimoniously explained by moisture and manure on the meadows in combination with seed mass and specific leaf area (SLA). In theGrazedGrassland data, species presence is parsimoniously explained by the grazing intensity and soil phosphorus in combination with theC:Nratio and flowering mode. Conclusions OurGLMMapproach can be used to identify which species traits and environmental variables best explain the species distribution, and which traits are significantly correlated with environmental variables. We argue that the method is better suited for providing an interpretable and predictive model than the fourth corner method andRLQ.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.026 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.008 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".