Extraction and high-throughput sequencing of oak heartwood DNA: Assessing the feasibility of genome-wide DNA methylation profiling
Bibliographic record
Abstract
Tree ring features are affected by environmental factors and therefore are the basis for dendrochronological studies to reconstruct past environmental conditions. Oak wood often provides the data for these studies because of the durability of oak heartwood and hence the availability of samples spanning long time periods of the distant past. Wood formation is regulated in part by epigenetic mechanisms such as DNA methylation. Studies of the methylation state of DNA preserved in oak heartwood thus could identify epigenetic tree ring features informing on past environmental conditions. In this study, we aimed to establish protocols for the extraction of DNA, the high-throughput sequencing of whole-genome DNA libraries (WGS) and the profiling of DNA methylation by whole-genome bisulfite sequencing (WGBS) for oak (Quercus robur) heartwood drill cores taken from the trunks of living standing trees spanning the AD 1776-2014 time period. Heartwood contains little DNA, and large amounts of phenolic compounds known to hinder the preparation of high-throughput sequencing libraries. Whole-genome and DNA methylome library preparation and sequencing consistently failed for oak heartwood samples more than 100 and 50 years of age, respectively. DNA fragmentation increased with sample age and was exacerbated by the additional bisulfite treatment step during methylome library preparation. Relative coverage of the non-repetitive portion of the oak genome was sparse. These results suggest that quantitative methylome studies of oak hardwood will likely be limited to relatively recent samples and will require a high sequencing depth to achieve sufficient genome coverage.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".