Analyzing the shape of observed trait distributions enables a data‐based moment closure of aggregate models
Bibliographic record
Abstract
Abstract The shape of trait distributions may inform about the selective forces that structure ecological communities. Here, we present a new moment‐based approach to classify the shape of observed biomass‐weighted trait distributions into normal, peaked, skewed, or bimodal that facilitates spatio‐temporal and cross‐system comparisons. Our observed phytoplankton trait distributions exhibited substantial variance and were mostly skewed or bimodal rather than normal. Additionally, mean, variance, skewness und kurtosis were strongly correlated. This is in conflict with trait‐based aggregate models that often assume normally distributed trait values and small variances. Given these discrepancies between our data and general model assumptions we used the observed trait distributions to test how well different aggregate models with first‐ or second‐order approximations and different types of moment closure predict the biomass, mean trait, and trait variance dynamics using weakly or moderately nonlinear fitness functions. For weakly non‐linear fitness functions aggregate models with a second‐order approximation and a data‐based moment closure that relied on the observed correlations between skewness and mean, and kurtosis and variance predicted biomass and often also mean trait changes fairly well and better than models with first‐order approximations or a normal‐based moment closure. In contrast, none of the models reflected the changes of the trait variances reliably. Aggregate model performance was often also poor for moderately nonlinear fitness functions. This questions a general applicability of the normal‐based approach, in particular for predicting variance dynamics determining the speed of trait changes and maintenance of biodiversity. We evaluate in detail how and why better approximations can be obtained.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".