Rarefaction and extrapolation of species richness using an area‐based Fisher's logseries
Bibliographic record
Abstract
Fisher's logseries is widely used to characterize species abundance pattern, and some previous studies used it to predict species richness. However, this model, derived from the negative binomial model, degenerates at the zero-abundance point (i.e., its probability mass fully concentrates at zero abundance, leading to an odd situation that no species can occur in the studied sample). Moreover, it is not directly related to the sampling area size. In this sense, the original Fisher's alpha (correspondingly, species richness) is incomparable among ecological communities with varying area sizes. To overcome these limitations, we developed a novel area-based logseries model that can account for the compounding effect of the sampling area. The new model can be used to conduct area-based rarefaction and extrapolation of species richness, with the advantage of accurately predicting species richness in a large region that has an area size being hundreds or thousands of times larger than that of a locally observed sample, provided that data follow the proposed model. The power of our proposed model has been validated by extensive numerical simulations and empirically tested through tree species richness extrapolation and interpolation in Brazilian Atlantic forests. Our parametric model is data parsimonious as it is still applicable when only the information on species number, community size, or the numbers of singleton and doubleton species in the local sample is available. Notably, in comparison with the original Fisher's method, our area-based model can provide asymptotically unbiased variance estimation (therefore correct 95% confidence interval) for species richness. In conclusion, the proposed area-based Fisher's logseries model can be of broad applications with clear and proper statistical background. Particularly, it is very suitable for being applied to hyperdiverse ecological assemblages in which nonparametric richness estimators were found to greatly underestimate species richness.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".