Dataset for "NanoVar: a Comprehensive Workflow for Structural Variant Detection to uncover the Genome's Hidden Patterns"
Bibliographic record
Abstract
Output Files for Long-Read Structural Variant and Repeat Analysis in Colorectal Cancer Samples (HRR698464, HRR698460, C586, C588) Description: This Zenodo dataset includes comprehensive output files generated during the application of a long-read sequencing analysis protocol for structural variant (SV) detection and repeat element characterization in colorectal cancer samples. The dataset is organized into two main directories: 1. HRR698464_MSI-H_TumorThis directory contains all primary output files generated from the analysis pipeline applied to the MSI-H tumor sample HRR698464 (also referred to as patient C586.T). Each subdirectory corresponds to a specific stage in the protocol: NanoPlot_outputOutput from Stage 1 – Quality assessment of raw reads using NanoPlot. SAMtools_outputBAM file processing outputs from Stage 2 – Alignment of long reads to the reference genome using SAMtools. NanoVar_outputOutput from Stage 3 – Structural variant calling using NanoVar. VCF_filtering_outputOutput from Stage 4 – Filtering of structural variants using SURVIVOR and BCFtools; includes the filtered VCF files. NanoINSight_outputOutput from Stage 5 – Characterization of repeat elements using NanoINSight. VEP_outputOutput from Stage 6 – Annotation of structural variants using Ensembl Variant Effect Predictor (VEP). 2. Additional_output_filesThis directory contains supplementary output files used for comparison and visualization in Figures 4–7 of the associated publication. These include: HRR698460.NanoPlot.report.htmlNanoPlot quality summary of a lower-quality tumor sample (HRR698460), used in Figure 4 for comparison with HRR698464. C586.N.nanovar.pass.vcfNanoVar VCF output for the matched normal sample of patient C586, used to filter somatic calls in Stage 4. C586.N.nanovar.pass.report.htmlNanoVar summary report of the normal sample of C586; used in Figure 5a. C588.N.nanovar.pass.vcfNanoVar VCF output of the MSS normal sample (C588) for comparison with the MSI-H patient (C586). C588.N.nanovar.pass.report.htmlNanoVar summary report of the MSS normal sample; used in Figure 5b. C588.T.nanovar.pass.vcfNanoVar VCF output of the MSS tumor sample (C588); used in comparative analyses with the MSI-H sample. C588.T.nanovar.pass.report.htmlNanoVar summary report of the MSS tumor sample; used in Figure 5b. MSS.tumor.unique.vcfVCF file of somatic SVs in the MSS sample, generated by comparing matched tumor and normal pairs. MSS.tumor.unique.RepeatMasker.tblRepeatMasker output annotating somatic insertions in the MSS tumor sample; used in Figure 6. MSS.tumor.unique.vep.htmlEnsembl VEP annotation report of somatic SVs in the MSS patient; used in Figures 7a and 7b. This dataset supports reproducibility and transparency of the protocol and offers a valuable resource for researchers interested in long-read-based SV detection, repeat annotation, and comparative cancer genomics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.068 | 0.060 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".