Operationalizing Quality Assurance for Clinical Illumina Somatic Next-Generation Sequencing Pipelines
Notice bibliographique
Résumé
Quality assurance (QA) is essential for precision oncology workflows, in particular in the clinical setting. However, because of numerous variations in laboratory and bioinformatics pipelines, QA practices remain non-standardized, are often ad hoc, and are lacking longitudinal tracking. A selected review of existing software was performed for quality control of Illumina next-generation sequencing data, focusing specifically on generalizable tools that can be integrated into any bioinformatics workflow to easily develop a QA workflow with longitudinal tracking. Although all implementations need to be integrated, validated, and iterated upon to suit individual operations, providing a base suite of options will enable better validation and use of QA in clinical somatic mutation testing for workflows using Illumina next-generation sequencing and beyond. Quality assurance (QA) is essential for precision oncology workflows, in particular in the clinical setting. However, because of numerous variations in laboratory and bioinformatics pipelines, QA practices remain non-standardized, are often ad hoc, and are lacking longitudinal tracking. A selected review of existing software was performed for quality control of Illumina next-generation sequencing data, focusing specifically on generalizable tools that can be integrated into any bioinformatics workflow to easily develop a QA workflow with longitudinal tracking. Although all implementations need to be integrated, validated, and iterated upon to suit individual operations, providing a base suite of options will enable better validation and use of QA in clinical somatic mutation testing for workflows using Illumina next-generation sequencing and beyond. Quality assurance (QA) is an essential part of clinical next-generation sequencing (NGS). To ensure reliable results, each step of a clinical NGS pipeline must be reviewed and validated for every sample in every run. QA ensures that each sample meets stringent quality standards and that quality remains consistent over time. Furthermore, appropriately designed QA processes can detect drifts in performance early, before they cause major issues. An ideal QA implementation would fulfill these goals. Although guidelines for QA and quality control (QC) for clinical NGS have been published, details on how laboratories might implement such guidelines are not available. NGS laboratories often build QA protocols from scratch, which can be laborious, time-consuming, noncomprehensive, and expensive. Moreover, different protocols in different laboratories are generally not standardized, with subtle differences in QA implementation that can produce different results. For example, differences in methods in library insert size calculations can produce a wide variance in values that can reduce the utility of the metric. Specific guidance for how to implement a QA system based on published guidelines1Roy S. Coldren C. Karunamurthy A. Kip N.S. Klee E.W. Lincoln S.E. Leon A. Pullambhatla A. Temple-Smolkin R.L. Voelkerding K.V. Wang C. Carter A.B. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: a joint recommendation of the Association for Molecular Pathology and the College of American Pathologists.J Mol Diagn. 2018; 20: 4-27Abstract Full Text Full Text PDF PubMed Scopus (279) Google Scholar are addressed. Specific tools required to calculate the list of per-sample and per-pipeline run metrics will be presented first. In the second part, a combination of MultiQC version 1.122Ewels P. Magnusson M. Lundin S. Käller M. MultiQC: summarize analysis results for multiple tools and samples in a single report.Bioinformatics. 2016; 32: 3047-3048Crossref PubMed Scopus (3012) Google Scholar and ChronQC version 1.0.23Tawari N.R. Seow J.J.W. Perumal D. Ow J.L. Ang S. Devasia A.G. Ng P.C. ChronQC: a quality control monitoring system for clinical next generation sequencing.Bioinformatics. 2018; 34: 1799-1800Crossref PubMed Scopus (1) Google Scholar will be demonstrated to display these metrics, as well as track them longitudinally to monitor process trends. Given the diversity of sample types, preparation methods, sequencing platforms, and bioinformatics tools, it is important to establish the scope and limitations of this article. The metrics recommendations are limited to short-read DNA sequencing on Illumina (San Diego, CA) platforms for tumor-only targeted panels. In addition, only metrics from Roy et al1Roy S. Coldren C. Karunamurthy A. Kip N.S. Klee E.W. Lincoln S.E. Leon A. Pullambhatla A. Temple-Smolkin R.L. Voelkerding K.V. Wang C. Carter A.B. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: a joint recommendation of the Association for Molecular Pathology and the College of American Pathologists.J Mol Diagn. 2018; 20: 4-27Abstract Full Text Full Text PDF PubMed Scopus (279) Google Scholar were considered and do not include metrics related to per-variant review metrics or metrics that cannot be inferred bioinformatically. In short, this article gives practical instructions on how to operationalize a QA protocol for clinical Illumina NGS sequencing. These recommendations can provide a start for clinical laboratories looking to build standardized, reproducible pipelines, and many of these recommendations can be exported to other sequencing workflows with modifications. Metrics were selected from the fourth table of the article by Roy et al,1Roy S. Coldren C. Karunamurthy A. Kip N.S. Klee E.W. Lincoln S.E. Leon A. Pullambhatla A. Temple-Smolkin R.L. Voelkerding K.V. Wang C. Carter A.B. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: a joint recommendation of the Association for Molecular Pathology and the College of American Pathologists.J Mol Diagn. 2018; 20: 4-27Abstract Full Text Full Text PDF PubMed Scopus (279) Google Scholar “Recommended Quality Metrics for Clinical Bioinformatics Pipelines,” and evaluated for ease of implementation in a clinical, tumor-only, small laboratory context with limited bioinformatics resources. Several quality metrics were excluded, such as number of germline single-nucleotide variants (SNVs), homozygous/heterozygous ratio, and DNA concentration, because these were either not relevant to somatic sequencing or referred to QC that would need to be applied before sequencing. Regarding DNA concentration, it is difficult to accurately determine concentration through bioinformatics alone and should be assessed before sequencing. As the focus of the article is operationalization of QA with minimal custom programming and manual software installs, bioinformatics software leveraged by the bcbio version 1.2.9 somatic small variant pipeline (https://doi.org/10.5281/zenodo.5781867, last accessed July 26, 2023) was investigated and given preference for which metrics are generated by software used by the pipeline. For each metric that could not be generated through bcbio-adjacent version 1.2.9 software, selected literature reviews were performed to investigate common bioinformatics software used in metrics generation. Results were evaluated for compatibility with reporting software MultiQC version 1.12, availability with Bioconda version 0.19.4,4Grüning B. Dale R. Sjödin A. Chapman B.A. Rowe J. Tomkins-Tinch C.H. Valieris R. Köster J. Biocanda TeamBioconda: sustainable and comprehensive software distribution for the life sciences.Nat Methods. 2018; 15: 475-476Crossref PubMed Scopus (446) Google Scholar integration with the bcbio version 1.2.9 somatic small variant pipeline, calculation method (if multiple programs were available), open-sourced, freely available, state of software maintenance, and ease of parsing output. The state of software maintenance was defined as having been updated within the past 3 years. In cases where the software was not updated in the past 3 years, it was investigated if there was a need for an update by reviewing posted issues in their associated repository, if available. Supplemental Code S1 provides the code for installation of all programs used. For sample-level metrics, software was sought to present metrics in a report format aggregated with other samples. For longitudinal graphs, software was evaluated using a wide range of criteria, including installation, integration with other software, reliability, automatability, user alerting, configurability, interactivity with core features for the generation of Levey-Jennings control charts,5Levey S. Jennings E.R. The use of control charts in the clinical laboratory.Am J Clin Pathol. 1950; 20: 1059-1066Crossref PubMed Scopus (0) Google Scholar easy uptake of delimited file formats, the application of Westgard rules,6Westgard J. Barry P. Cost-effective Quality Control: Managing the Quality and Productivity of Analytical Processes. AACC Press, Washington, DC1986Google Scholar and software license status. OmnomicsQ version 1.0.42.0 (Euformatics, Espoo, Finland) was used as a feature comparator. A review of features can be found in Supplemental Table S1. Selected literature reviews were conducted within PubMed, Google Scholar, and Google. Searches were conducted from October 2019 to April 2020. Two sets of searches were conducted. The first was the search for metric generation methods that were not already fulfilled by the Genomic Analysis Toolkit (GATK) version 4.2.6.1 (Broad Institute, Cambridge, MA)7Van der Auwera G.A. Carneiro M.O. Hartl C. Poplin R. Del Angel G. Levy-Moonshine A. Jordan T. Shakir K. Roazen D. Thibault J. Banks E. Garimella K.V. Altshuler D. Gabriel S. DePristo M.A. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit practices PubMed Scopus Google Scholar or a base bcbio version 1.2.9 For of the on search used of the metric metrics, and For of that are the search used were metrics, and of the for not all bioinformatics to be published, in the Genome were for for A second of searches were for software that could display metrics for and longitudinal that were used in all were quality Westgard and In to the search results, related software for major QC programs was of these programs are version and for version J. M. results into a and quality control PubMed Scopus Google Scholar searches for sets were conducted using of the search in this and the search such as for The data sets generated the are in the last accessed July 26, data used are available. the review of the metrics based on the criteria, recommendations for in Table have been der Auwera G.A. Carneiro M.O. Hartl C. Poplin R. Del Angel G. Levy-Moonshine A. Jordan T. Shakir K. Roazen D. Thibault J. Banks E. Garimella K.V. Altshuler D. Gabriel S. DePristo M.A. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit practices PubMed Scopus Google G. An and analysis for variant and from DNA PubMed Scopus Google Scholar, K. A. quality control for sequencing 2016; 32: PubMed Scopus Google Scholar, B. A. T. J. G. G. R. Genome format and PubMed Scopus Google Scholar, P. A. Wang M. T. Wang A for and the of single in the of PubMed Scopus Google Scholar are Bioconda version using version which is a of the software and required for to metrics from the report by the programs have been with a focus on implementations the need for many will have on the and of the of to Quality and version der Auwera G.A. Carneiro M.O. Hartl C. Poplin R. Del Angel G. Levy-Moonshine A. Jordan T. Shakir K. Roazen D. Thibault J. Banks E. Garimella K.V. Altshuler D. Gabriel S. DePristo M.A. From FastQ data to high confidence variant calls: the Genome Analysis Toolkit practices PubMed Scopus Google version version version version G. An and analysis for variant and from DNA PubMed Scopus Google version K. A. quality control for sequencing 2016; 32: PubMed Scopus Google version B. A. T. J. G. G. R. Genome format and PubMed Scopus Google version P. A. Wang M. T. Wang A for and the of single in the of PubMed Scopus Google performed through Bioconda version Genomic Analysis table in a performed through Bioconda version Genomic Analysis of metrics, in Table is as or These metrics can be the or Although monitoring first the is there are cases where individual can in performance because of for example, with the or of the on the Illumina The and system of each Illumina can and can produce different metrics can be and through the Analysis can be by metrics through the library version this will parsing and calculation of the output. on these can be found in Supplemental Table of Metrics with Bioinformatics and of a of the of all table in a is as the number of an Illumina which is an for sequencing output. can quality high a can with the of Analysis software to accurately the from individual of data in Illumina sequencing can be by Scopus Google Scholar could is not limited of samples or DNA in the confidence of the base can be through the version The number of a can be a for sequencing and quality of the A will as for the metric is by wide range of such as sample sample and can be reviewed with the as For example, a with a high number of and a to high The number of can be through the version The of the of all the Illumina quality of the Although this metric is to the in the it the could issues related to the quality of the and high For example, a library could the metric base quality the of metric can be using the version that quality can be Diego, Scholar on the and version software used. the on the A high of is to be metric can be by the the and version Although the using a metric for performance in S. Coldren C. Karunamurthy A. Kip N.S. Klee E.W. Lincoln S.E. Leon A. Pullambhatla A. Temple-Smolkin R.L. Voelkerding K.V. Wang C. Carter A.B. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: a joint recommendation of the Association for Molecular Pathology and the College of American Pathologists.J Mol Diagn. 2018; 20: 4-27Abstract Full Text Full Text PDF PubMed Scopus (279) Google Scholar can be as there are often cases where samples do not Although every sequencing and will have different library it is for to be to for longitudinal of the metric. can be through the version can be per-sample metric. Metrics within this will on the sample and are generated or not the metric is through an run of bcbio version 1.2.9 is in Table of Metrics with the Bioinformatics and for version 1.2.9 of of targeted with a of the on insert Genomic Analysis single-nucleotide bcbio version 1.2.9 table in a Genomic Analysis single-nucleotide of the of within a metric is to ensure the user that the sample library a high over to ensure the to variants the variant The results of this metric are by many such as the quality of library sequencing DNA concentration, and which should be reviewed in cases of in this metric. However, this metric is to detect individual metric can be through the within version (Broad Institute, Cambridge, of targeted with a the of a within metric is to monitor where or in can individual or that were in the of metric. is because for targeted can and a high of there could be targeted with that would not be for variant An of variance can be in Supplemental S1. metric can be through the version 4.2.6.1 with the to the The will be on the or clinical and sequencing sample and of As an for variant with quality and variant of could with E.R. S. A. of in clinical Mol Diagn. Full Text Full Text PDF PubMed Scopus Google Scholar for this metric and they their for The is to a that provides an that will detect of a will could issues or such as or high a would an of could be used as of the on the of a within a metric as and However, it is performed the sample and on the can be using the version G. An and analysis for variant and from DNA PubMed Scopus Google Scholar and parsing the output. as the metrics, the should into all minimal used variant For example, version D. E.R. somatic mutation and number in by PubMed Scopus Google Scholar a of by which is for Illumina sequencing for clinical somatic variant the of within the sample of and As the is not to samples within an this metric as a metric used in with for sample preparation and quality issues. could issues in such as of and of high or C. T. Wang J. M. G. Wang J. Wang J. of PubMed Scopus Google R. R. G. M. of DNA sequencing PubMed Scopus Google Scholar metric can be through the in the version K. A. quality control for sequencing 2016; 32: PubMed Scopus Google Scholar In addition, this metric can be through version generation. insert size in the context of Illumina sequencing is a of the the DNA within the which the size process can be by other size can issues with and sample C. C. P. M. J. G. J. A high of is to of J Pathol. Full Text Full Text PDF PubMed Google Scholar The metric is often only in as the metric is generated in the NGS pipeline size distribution is often of bioinformatics through manual review of generated by for C. J. P. M. and quality in next-generation 2016; PubMed Scopus Google Scholar this metric can be from However, to a can produce results in As variants and can to insert this can the and In addition, different software tools have different methods for results. version the distribution of insert to from the with a of version B. A. T. J. G. G. R. Genome format and PubMed Scopus Google P. J. J. M.O. A. T. of and Scopus Google Scholar only insert for the calculation a which by is version to variations of results. For a first this metric using version or version is is a of the of from can from multiple including from and D. M. T. C. C. A. and in Illumina sequencing PubMed Scopus Google Scholar To and this the of and use must be considered metric can be of of DNA high which could to Results can be generated from version using the is a confidence for sample can be inferred on the and if and as a for the have to be for each laboratory as this calculation would be by through and that should be as this would in can samples with of R. of in to PubMed Scopus Google Scholar, C. M. M. Magnusson D. M. E. E. is associated with of PubMed Scopus Google Scholar, M. J. B. J. E. J. J. M. in the of in of J PubMed Scopus Google Scholar The can be generated by the of the version 4.2.6.1 for and and is a of the number of the number of within a can be of the of and other Given that this metric can using it in with other metrics to on the quality of sample is can be by parsing the of version P. A. Wang M. T. Wang A for and the of single in the of PubMed Scopus Google Scholar or from the variant format as the of to with an and for the and 3 for the this will by preparation M.A. Banks E. Poplin R. Garimella K.V. Hartl C. Angel M. M.A. M. A. K. Gabriel Altshuler D. A for and using next-generation DNA sequencing PubMed Scopus Google Scholar metric can be in a germline context as from the a However, in a somatic custom there can be variations in the because of the mutation of the In addition, the can be on the mutation and other and will need to be an individual sample can be by the of version 4.2.6.1 and parsing the or from the variant format of that are is as and metric in a clinical guidance for through variant within the metric is of the bioinformatics pipeline often by a before sample However, this can be and because of of the can be inferred using programs such as K. J. M.A. analysis of the and of Scopus (0) Google Scholar T. M. number and mutation from sequencing Full Text Full Text PDF PubMed Scopus Google Scholar and P. J. E. C. K. S. A. P. M. M. S. E. of PubMed Scopus Google Scholar However, given the of types, of somatic and sample preparation on a recommendation is not and should with different programs to the in their methods can be with a on the of the Although a to ensure quality assurance is a of metrics, the is that these metrics should be of types, and monitoring can be through sample-level that of all of performance to samples for Although monitoring is well to detect from subtle such as would be the to a to is a major is as if the through and consistent is if any are based on a such by the was a in within a maintenance of of the and from a particular over a of samples to a cause could have the performance A second was a in an that the in used for DNA in a for a targeted in high that of of which can be in Supplemental S1. the of longitudinal through the application of Westgard and Levey-Jennings as been applied in clinical laboratories for S. Jennings E.R. The use of control charts in the clinical laboratory.Am J Clin Pathol. 1950; 20: 1059-1066Crossref PubMed Scopus (0) Google J. of QC for clinical Scopus (0) Google Scholar, A. Quality control for the of somatic in clinical by next-generation Pathol. PubMed Google Scholar, J. E. J. M. E. K. PubMed Scopus (0) Google Scholar Levey-Jennings charts are a of control for quality The metric evaluated is the the data are on the of the are of with the on the Westgard to a of to where an run should be J. Barry P. Cost-effective Quality Control: Managing the Quality and Productivity of Analytical Processes. AACC Press, Washington, DC1986Google Scholar of Westgard are not limited a that in the run that have differences and on the of the can issues that might be only in a is through the combination of and longitudinal monitoring that quality assurance is For MultiQC version is MultiQC version common bioinformatics QC and MultiQC version and all metrics by report it will metrics is in this article. MultiQC for custom The first with the metrics and as for their to user review A MultiQC version report can be in with the report the MultiQC version last accessed July 26, The report features from the version version version and other tools, and can from a of including custom generated by the with is recommendation was to be with MultiQC version with minimal The report a sample with high In the version generated of the version the of the of how each metric is generated and the need to develop protocols which should be Supplemental Table which report file is used for each metric. For the a QA monitor should the and it for a cause and that if this were to be the it could be the longitudinal review if the Westgard was MultiQC version 1.12, a user can the performance of samples in the context of an run and multiple metrics to for of relevant issues. The generated user with features such as custom and custom sample or through or An of custom sample can be in Supplemental MultiQC version not implement for within can be through Levey-Jennings For longitudinal for the were provide a of the data, data to the the to and implementation of Westgard For a list of Supplemental Table S1. From the selected a that fulfilled all features was not However, ChronQC version fulfilled core of longitudinal such as Levey-Jennings of and the implementation of the Westgard Westgard have not been a of longitudinal monitoring for the metric with a to The a Westgard in 2020. A QA monitor should the and from that for In this example, the cause was to a concentration of samples for that related to of samples. The number of samples with DNA was related to a validation of the pipeline, where only samples were available. was this the of the QA quality assurance of a clinical NGS pipeline the generation of a of metrics and the review of the metrics, for a of Metrics need to the and performance of sample library NGS and bioinformatics reviews need to performance in and longitudinal of are presented for the generation of metrics in a somatic Illumina short-read sequencing context through a of tools and through Bioconda version with tools within bcbio version However, these can be short-read sequencing. are to and the metrics from report as in Supplemental Table The are designed to minimal user and of with minimal installation and of The use only common and the for These are not to be the for an individual provide a are to review the report to their methods and metrics for quality A pipeline using which the use of these is last accessed July 26, metrics to in the of the bioinformatics pipeline to be of all the metric generation programs as many of the will the a a is a these programs can be run and from the bioinformatics pipeline. In addition, selected for metrics should be to that they are to with a minimal the metrics are they are presented in an format through MultiQC version review and ChronQC version longitudinal Levey-Jennings control these this article the first to the operationalization of a quality assurance for clinical somatic Illumina NGS looking to build reproducible that with to detect a of of the focus on Illumina the metrics, the tools can be applied to other sequencing platforms and including sequencing. In QA ensures that each sample meets stringent quality standards and that quality remains consistent over time. these recommendations and appropriately designed QA processes can detect drifts in performance early, before they cause major issues. the Genome quality control and for and and the and bioinformatics data and the and designed and performed and and the and the Supplemental of a MultiQC version report with custom the is the for the MultiQC The is and the search been will sample and in the table and the the metrics for the of insert size of of variants and The of the is a based on results from a for with Supplemental Table S1 with Supplemental Table with Supplemental Code S1
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,011 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».