Plant DNA‐barcode library and community phylogeny for a semi‐arid East African savanna
Bibliographic record
Abstract
Applications of DNA barcoding include identifying species, inferring ecological and evolutionary relationships between species, and DNA metabarcoding. These applications require reference libraries that are not yet available for many taxa and geographic regions. We collected, identified, and vouchered plant specimens from Mpala Research Center in Laikipia, Kenya, to develop an extensive DNA-barcode library for a savanna ecosystem in equatorial East Africa. We amassed up to five DNA barcode markers (rbcL, matK, trnL-F, trnH-psbA, and ITS) for 1,781 specimens representing up to 460 species (~92% of the known flora), increasing the number of plant DNA barcode records for Africa by ~9%. We evaluated the ability of these markers, singly and in combination, to delimit species by calculating intra- and interspecific genetic distances. We further estimated a plant community phylogeny and demonstrated its utility by testing if evolutionary relatedness could predict the tendency of members of the Mpala plant community to have or lack "barcode gaps", defined as disparities between the maximum intra- and minimum interspecific genetic distances. We found barcode gaps for 72%-89% of taxa depending on the marker or markers used. With the exception of the markers rbcL and ITS, we found that evolutionary relatedness was an important predictor of barcode-gap presence or absence for all of the markers in combination and for matK, trnL-F, and trnH-psbA individually. This plant DNA barcode library and community phylogeny will be a valuable resource for future investigations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".