Do plant traits retrieved from a database accurately predict on‐site measurements?
Bibliographic record
Abstract
Summary Trait‐based approaches are increasingly used to obtain an insight into the functional aspects of plant communities. Since measuring traits can be time‐consuming, large international databases of plant traits are being compiled to share the effort. From these databases, average trait values are often extracted per species by averaging trait values of individuals over multiple populations and habitats. However, the accuracy of such aggregated information from regional databases as a surrogate for on‐site measurements has seldom been tested. For the local species pool (aggregated at the habitat‐level) and the plant communities on the plots (aggregated at the community‐level), we quantified how accurately trait values for each species measured at the plot (plot scale) and those averaged per species and site (site scale) can be estimated from those retrieved from a North‐west‐European trait database. We analysed three widely used plant traits, canopy height ( CH ), leaf dry matter content ( LDMC ) and specific leaf area ( SLA ), of species occurring in a wet meadow and a salt marsh. Database values more accurately predicted traits aggregated at the habitat‐level than those aggregated at the community‐level. In addition, traits with lower plasticity, such as LDMC , were more accurately predicted by database values. The performance of database values also depended upon the habitat studied, for example, habitat‐level SLA values were accurately predicted by database values in the wet meadow but inaccurately predicted in the salt marsh. Synthesis . This study reveals that the accuracy of traits retrieved from a database depends on the level of aggregation (lower at community‐level), the trait (lower in plastic traits) and the habitat type (lower in extreme habitats). For studies focussing on processes mainly acting at the site scale (e.g. trait–environment relationships), traits retrieved from a regional database and filtered according to habitat will probably lead to good results. Whereas studying processes acting at the plot scale (e.g. niche partitioning), requires the additional effort of measuring traits on‐site.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".