Bibliographic record
Abstract
This release modernizes nf-ionampliseq to DSL-2 standards with enhanced BLAST analysis capabilities, improved containerization, and better resource management. Key additions include BLAST database integration, enhanced variant calling options, and comprehensive process version tracking. Added [feat] BLASTN process to run blastn with consensus sequences against user-specifed DB (--blast_db). [feat] BLASTN_COVERAGE process and blast_coverage.py for summarizing BLAST results and generating MultiQC tables. [feat] fill-tags plugin for BCFTOOLS_FILTER and better filtering of variants for consensus sequence construction. [feat] SEQTK_SUBSEQ process for extracting reference sequences based on Mash screen results. [feat] CONSENSUS_MULTIQC process for enhanced consensus sequence reporting in MultiQC. [feat] Enhanced container support for Docker, Singularity, Apptainer, and Podman. [feat] Version tracking for all processes with versions.yml files. [config] comprehensive DSL2 module configuration in conf/modules.config. [config] new configuration parameters: blast_db, trim_primers, output_unmapped_reads. [config] variant calling options: minor_allele_fraction, major_allele_fraction, low_coverage. [config] enhanced container profiles for multiple container engines. [config] proper resource management with check_max function for memory, CPU, and time limits. [config] timestamped execution reports and timeline files in nextflow.config. Changes [docs] Comprehensive documentation overhaul with detailed parameter descriptions and external tool references [cleanup] Removed legacy plotting module and environment.yml file for simplified architecture. Use wgscovplot instead for better and interactive plotting of coverage stats and variants. [cleanup] Simplified TMAP process configuration using dynamic arguments parameter. [update] Repository and container registry references updated to CFIA-NCFAD organization [feat] Complete workflow modernization to DSL2 standards with proper process separation. [feat] Added new Docker build workflow in .github/workflows/docker.yml for improved container publishing and CI separation. [feat] Added support for building and using custom Docker images with updated samtools and TMAP/TVC binaries. [feat] Enhanced process version tracking: all major processes now emit versions.yml with tool versions for reproducibility. [feat] Improved error handling and retry strategies for all processes. [update] MultiQC process updated to support new BLAST coverage MultiQC table. [update] Bump version of workflow and dependencies for 2.0.0 release. [update] Repository renamed from peterk87/nf-ionampliseq to CFIA-NCFAD/nf-ionampliseq. [update] Nextflow version requirement updated to !>=22.10.1. [update] All processes now use proper container definitions with conda and container specifications. [update] Process resource labeling standardized (process_low, process_medium, process_high). [update] Input/output patterns standardized across all processes with proper emit declarations. [update] TMAP and TVC processes enhanced with better parameter handling and version tracking. [update] TMAP and TVC processes now use new Docker image ghcr.io/cfia-ncfad/nf-ionampliseq:2.0.0 with updated samtools and runtime dependencies. [update] MASH screen workflow restructured with separate sketching and screening processes. [update] FastQC process enhanced with memory optimization and version tracking. [update] Samtools processes standardized with consistent version tracking and output handling. [update] Mosdepth process optimized with improved output handling and version tracking. [update] Edlib processes enhanced with container support and version tracking. [update] Sample sheet processing improved with better error handling and validation. [docs] Documentation and help text for BLAST coverage analysis improved. [docs] Updated documentation to reflect new container build and usage instructions. [ci] Updated GitHub Actions CI tests. [ci] Moved Docker container build to separate workflow for improved CI/CD. [ci] CI now tests with multiple Nextflow versions 22.10.1 and latest stable (25.04.6 currently). Fixed [fix] Container compatibility issues across different container engines. [fix] Process resource allocation and memory management. [fix] Version tracking consistency across all processes. [fix] Input/output file handling and validation. [fix] MultiQC report generation and custom content integration. Dependencies [deps] Updated to Nextflow !>=22.10.1. [deps] Enhanced container support with multiple engine options. [deps] Standardized conda environment specifications across all processes. [deps] Added better version tracking for all bioinformatics tools. What's Changed DSL2 modernization, BLAST analysis, and enhanced containerization by @peterk87 in https://github.com/CFIA-NCFAD/nf-ionampliseq/pull/1 Release 2.0.0 by @peterk87 in https://github.com/CFIA-NCFAD/nf-ionampliseq/pull/2 New Contributors @peterk87 made their first contribution in https://github.com/CFIA-NCFAD/nf-ionampliseq/pull/1 Full Changelog: https://github.com/CFIA-NCFAD/nf-ionampliseq/compare/1.0.1...2.0.0
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.013 |
| Meta-epidemiology (narrow) | 0.003 | 0.004 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.006 | 0.004 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.168 | 0.249 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".