Combining Off‐flow, a Nextflow‐coded program, and whole genome sequencing reveals unintended genetic variation in CRISPR/Cas-edited iPSCs
Bibliographic record
Abstract
Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-Cas nucleases and human induced pluripotent stem cell (iPSC) technology can reveal deep insight into the genetic and molecular bases of human biology and disease. Undesired editing outcomes, both on-target (at the edited locus) and off-target (at other genomic loci) hinder the application of CRISPR-Cas nucleases. We developed Off-flow, a Nextflow-coded bioinformatic workflow that takes a specific guide sequence and Cas protein input to call four separate off-target prediction programs (CHOPCHOP, Cas-Offinder, CRISPRitz, CRISPR-Offinder) to output a comprehensive list of predicted off-target sites. We applied it to whole genome sequencing (WGS) data to investigate the occurrence of unintended effects in human iPSCs that underwent repair or insertion of disease-related variants by homology-directed repair. Off-flow identified a 3-base-pair-substitution and a mono-allelic genomic deletion at the target loci, KCNQ2 , in 2 clones. Unbiased WGS analysis further identified off-target missense variants and a mono-allelic genomic deletion at the targeted locus, GNAQ, in 10 clones. On-target substitution and deletions had escaped standard PCR and Sanger sequencing analysis, while missense variants at other genomic loci were not detected by Off-flow. We used these results to filter out iPSC clones for subsequent functional experiments. Off-flow, which we make publicly available, works for human and mouse genomes currently and can be adapted for other genomes. Off-flow and WGS analysis can improve the integrity of studies using CRISPR/Cas-edited cells and animal models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".