MétaCan
Menu
Back to cohort
Record W3004211264 · doi:10.17504/protocols.io.7euhjew

Rocky Mountain adventures in Genomic DNA sample preparation, ligation protocol optimisation / simplification and Ultra long read generation v1

2019· preprint· en· W3004211264 on OpenAlexaff
John R. Tyson

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicMolecular Biology Techniques and Applications
Canadian institutionsUniversity of British Columbia Hospital
Fundersnot available
KeywordsWorkflowDNA sequencinggenomic DNAComputer scienceProtocol (science)LigationDNASample (material)Computational biologyBiologyChemistryGeneticsMedicineMolecular biologyChromatographyDatabase

Abstract

fetched live from OpenAlex

We have been playing quite a bit with native genomic DNA sequencing for a project we have running on Epigenetic modifications of Rat models of neurological disease. While we are still rapidly iterating on tweaks etc. to some new protocols around increasing our production efficiency of ultra-long reads (100kb+) using LSK109 ligation kit components and PEG/NaCl mediated DNA precipitation, we are settling into a workflow now that I think is worth sharing and has been generating very good results for us. That said remember this is a little “wild west” and perhaps for the more adventurous at this stage, use at your own risk etc. etc. and perhaps help with some tweaks of your own ;o). I’ll say now that I have been optimizing this on a very easy Rat cell line grown in culture and realise some people don’t have this luxury, however I think this is a very good jumping off point and allowed us to iterate on a consistent starting material and will myself be venturing off into tissue extractions now. There are currently three main areas to highlight and detail changes within a sample to sequence workflow that we think has allowed us to push the efficiency and yields of sequence data out. Firstly DNA extractions using Phenol/Chloroform and spooling out of HWM DNA to generate the starting DNA sample. Secondly input material shearing and modifications to the LSK109 ligation protocol that have yielded high efficiency libraries, and production of greater ultra-long reads numbers. Lastly the use of flowcell refueling and DNAseI clearing of the flowcell surface to remove “blocking” DNA and allow fresh library addition after a surface “reset”. I’m detailing all we have here is an attempt to allow people to fully understand what is going on as much as possible and so we can all use this information to make intelligent changes to the protocols etc. if you are so inclined. Commercial “black box kits” and solutions while great for defined on mass purpose and reproducibility make this process more opaque and often overly complex and expensive. A good example of this is the addition of just NaCl after the ligation reaction to precipitate the adapted DNA showing how knowing what you have provides simplification and economic sense for the end user in some situations, but more on that later…… Final thoughts…. One of my interests around nanopore sequencing and moving sequencing back to small labs and beyond with the MinION device is also mitigating the sample preparation costs while not sacrificing performance. Using cheap needles, Polyethylene Glycol, salt and old school molecular biology techniques we are seeing uncompromising performance for both yield and read lengths that I think even the “big” guys will use :o)). I will reproduce this post and methods at www.longreadclub.org soon so they will be easier to find and develop going forward, and we will try and provide protocols on www.protocols.io as well. Things on the to do list include: Further optimisation of PEG/salt parameters for size selection / short read elimination at different stages of library preparation. Refinement of shearing to provide increased 100kb+ reads from HMW DNA Look at replacing Phenol/Chloroform with salting out in initial HMW Genomic DNA preparation. Who knew that PEG and salt would be a route to cheap ultra-long reads, we all just need to keep tweaking away in an open fashion and who knows where we can get. Anyway that’s enough from me for now, happy Nanopore adventuring!

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.298
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.017
GPT teacher head0.315
Teacher spread0.299 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2019
Admission routes1
Has abstractyes

Explore more

Same topicMolecular Biology Techniques and ApplicationsFrench-language works237,207