MétaCan
Menu
← Back to cohort
Record W6893302055 · doi:10.5281/zenodo.16583523

Dataset for "NanoVar: a Comprehensive Workflow for Structural Variant Detection to uncover the Genome's Hidden Patterns"

2025· dataset· en· W6893302055 on OpenAlexaff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2025
Typedataset
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Rare Diseases
Canadian institutionsMemorial University of Newfoundland
Fundersnot available
KeywordsDirectoryPipeline (software)WorkflowSample (material)AnnotationStage (stratigraphy)Filter (signal processing)

Abstract

fetched live from OpenAlex

Output Files for Long-Read Structural Variant and Repeat Analysis in Colorectal Cancer Samples (HRR698464, HRR698460, C586, C588) Description: This Zenodo dataset includes comprehensive output files generated during the application of a long-read sequencing analysis protocol for structural variant (SV) detection and repeat element characterization in colorectal cancer samples. The dataset is organized into two main directories: 1. HRR698464_MSI-H_TumorThis directory contains all primary output files generated from the analysis pipeline applied to the MSI-H tumor sample HRR698464 (also referred to as patient C586.T). Each subdirectory corresponds to a specific stage in the protocol: NanoPlot_outputOutput from Stage 1 – Quality assessment of raw reads using NanoPlot. SAMtools_outputBAM file processing outputs from Stage 2 – Alignment of long reads to the reference genome using SAMtools. NanoVar_outputOutput from Stage 3 – Structural variant calling using NanoVar. VCF_filtering_outputOutput from Stage 4 – Filtering of structural variants using SURVIVOR and BCFtools; includes the filtered VCF files. NanoINSight_outputOutput from Stage 5 – Characterization of repeat elements using NanoINSight. VEP_outputOutput from Stage 6 – Annotation of structural variants using Ensembl Variant Effect Predictor (VEP). 2. Additional_output_filesThis directory contains supplementary output files used for comparison and visualization in Figures 4–7 of the associated publication. These include: HRR698460.NanoPlot.report.htmlNanoPlot quality summary of a lower-quality tumor sample (HRR698460), used in Figure 4 for comparison with HRR698464. C586.N.nanovar.pass.vcfNanoVar VCF output for the matched normal sample of patient C586, used to filter somatic calls in Stage 4. C586.N.nanovar.pass.report.htmlNanoVar summary report of the normal sample of C586; used in Figure 5a. C588.N.nanovar.pass.vcfNanoVar VCF output of the MSS normal sample (C588) for comparison with the MSI-H patient (C586). C588.N.nanovar.pass.report.htmlNanoVar summary report of the MSS normal sample; used in Figure 5b. C588.T.nanovar.pass.vcfNanoVar VCF output of the MSS tumor sample (C588); used in comparative analyses with the MSI-H sample. C588.T.nanovar.pass.report.htmlNanoVar summary report of the MSS tumor sample; used in Figure 5b. MSS.tumor.unique.vcfVCF file of somatic SVs in the MSS sample, generated by comparing matched tumor and normal pairs. MSS.tumor.unique.RepeatMasker.tblRepeatMasker output annotating somatic insertions in the MSS tumor sample; used in Figure 6. MSS.tumor.unique.vep.htmlEnsembl VEP annotation report of somatic SVs in the MSS patient; used in Figures 7a and 7b. This dataset supports reproducibility and transparency of the protocol and offers a valuable resource for researchers interested in long-read-based SV detection, repeat annotation, and comparative cancer genomics.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.068
Threshold uncertainty score0.226

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.005
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0020.003
Science and technology studies0.0010.000
Scholarly communication0.0020.001
Open science0.0030.002
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0680.060

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.257
Teacher spread0.239 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)→Same topicGenomics and Rare Diseases→French-language works237,207→