DNA metabarcoding of saproxylic beetles - Streamlining species identification for large-scale forest biomonitoring
Bibliographic record
Abstract
Forest ecosystems host most of the terrestrial biodiversity on Earth. Climate change scenarios predict an increase in the intensity and frequency of severe summer droughts, high temperatures, and infestations of pathogens and insects, causing high mortality of some keystone tree species. These changes will affect forestry policies and practices, strongly impacting biodiversity. Understanding the responses of biodiversity to forest decline is therefore essential to developing new climate-smart management options. Biomonitoring of forest insects relies on techniques involving laborious and expensive sampling procedures. For instance, the study of indicators such as saproxylic beetles is strongly impeded by their high abun- dance and diversity, and by the deficit in taxonomists able to identify them. Here, we propose and test the use of metabarcoding for bulk samples of saproxylic beetles, in combination with the assembly of a relevant barcode reference library, as a mean to streamline identification. Results: Using a set of three primer pairs targeting short fragments within the cytochrome c oxidase subunit I (COI) barcode, we analyzed through metabarcoding a set of 32 bulk samples of saproxylic beetles collected in France, containing hundreds of specimens that were all initially counted and identified using morphology. To test the efficiency of non-destructive analyses, we also sequenced libraries of amplicons directly obtained from the ethanol used for preserving the samples. Identifying the resulting reads with a newly assembled barcode library, we successfully recovered most species present in each of the samples. Furthermore, our samples were selected to take into account a variety of conditions and parameters possibly affecting the results (species diversity, relative abundance and biomass, sampling medium, and preservation method). Significance: The use of DNA metabarcoding to monitor forest biodiversity can significantly improve our capacity to measure, understand, and anticipate the impact of global changes on forests, thus enhancing conservation strategies and the sustainability of silvicultural practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".