An assay validation framework to compare and evaluate targeted environmental DNA assays for routine species monitoring
Bibliographic record
Abstract
Environmental DNA (eDNA) analysis utilises trace DNA released by organisms into their environment for species detection and is revolutionising non ‐ invasive species and biodiversity monitoring. However, this technology requires rigorous validation along the whole workflow – from field sampling to statistical analysis – to ensure appropriate and meaningful interpretation of results. Targeted eDNA assays are often validated within a specific system and with particular aims, but without fulfilling predefined criteria. Consequently, their applicability beyond initial development often remains undetermined. Additionally, there tends to be poor understanding of the uncertainties and limitations associated with already published assays and thus potentially inappropriate interpretation of the results they produce. The lack of a “gold standard” limits the incorporation of targeted eDNA assays into species monitoring and policy making by end-users and is therefore key for the future implementation of eDNA-based surveys. Here, we present a framework (https://edna-validation.com/) and user-friendly criteria for the classification of assays, which is based on previous validation efforts. A 5 ‐ level assay validation scale (“incomplete” to “operational”) was defined by reviewing the current eDNA literature and conducting a meta-analysis on sampling, laboratory practices, detection limits, and detection probabilities. The so far published single species eDNA assays were reviewed for their performance in this new framework and we identified steps within the validation process that often remain untouched. Finally, we provide guidance for end ‐ users as to which criteria are most important for validation and suggest how results obtained from assays at different levels of the validation scale should be interpreted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".