Bibliographic record
Abstract
An ever-growing catalog of human variants is hosted in the ClinVar database. In this database, submissions on a variant are combined into a multisubmitter record; and in the case of discordance in variant classification between submitters, the record is labeled as conflicting. The current study used ClinVar data to identify characteristics that would make variants more likely to be associated with the conflict class of variants. Furthermore, the Extreme Gradient Boosting algorithm was used to train classifier models to provide prediction of classification discordance for single submission variants in ClinVar database. Population allele frequency, the gene harboring the variant, variant type, consequence on protein, variant deleteriousness score, first submitter identity, and submission count were associated with conflict in variant classification. Using such features, the optimized classifier showed accuracy on the test set of 88% with the weighted average of precision, recall, and f1-score of 0.84, 0.88, and 0.85, respectively. There were pronounced associations between variant classification discordance and allele frequency, gene type, and the identity of the first submitter. The study provides the predicted discordance status for single-submitter variants deposited in ClinVar. This approach can be used to assess whether single-submitter variants are likely to be supported, or in conflict with, future entries; this knowledge may help laboratories with clinical variant assessment. An ever-growing catalog of human variants is hosted in the ClinVar database. In this database, submissions on a variant are combined into a multisubmitter record; and in the case of discordance in variant classification between submitters, the record is labeled as conflicting. The current study used ClinVar data to identify characteristics that would make variants more likely to be associated with the conflict class of variants. Furthermore, the Extreme Gradient Boosting algorithm was used to train classifier models to provide prediction of classification discordance for single submission variants in ClinVar database. Population allele frequency, the gene harboring the variant, variant type, consequence on protein, variant deleteriousness score, first submitter identity, and submission count were associated with conflict in variant classification. Using such features, the optimized classifier showed accuracy on the test set of 88% with the weighted average of precision, recall, and f1-score of 0.84, 0.88, and 0.85, respectively. There were pronounced associations between variant classification discordance and allele frequency, gene type, and the identity of the first submitter. The study provides the predicted discordance status for single-submitter variants deposited in ClinVar. This approach can be used to assess whether single-submitter variants are likely to be supported, or in conflict with, future entries; this knowledge may help laboratories with clinical variant assessment. Determining the effect of variants on human health is crucial for modern clinical molecular diagnostics. Because of technological advancements and evolving commercial landscapes, molecular genetic diagnostic approaches have become more accessible to clinicians. Genetic laboratories now offer services to more clinicians and genotype more patients for more genes than ever before. However, an increase in test ordering and reporting is likely to lead to disagreements between laboratories over the clinical interpretation of certain variants.1Yang S. Lincoln S.E. Kobayashi Y. Nykamp K. Nussbaum R.L. Topper S. Sources of discordance among germ-line variant classifications in ClinVar.Genet Med. 2017; 19: 1118-1126Abstract Full Text Full Text PDF PubMed Scopus (75) Google Scholar,2Amendola L.M. Muenzen K. Biesecker L.G. Bowling K.M. Cooper G.M. Dorschner M.O. Driscoll C. Foreman A.K.M. Golden-Grant K. Greally J.M. Hindorff L. Kanavy D. Jobanputra V. Johnston J.J. Kenny E.E. McNulty S. Murali P. Ou J. Powell B.C. Rehm H.L. Rolf B. Roman T.S. Van Ziffle J. Guha S. Abhyankar A. Crosslin D. Venner E. Yuan B. Zouk H. Sequencing Diagnostic Yield working group; Jarvik GP. Variant classification concordance using the ACMG-AMP variant interpretation guidelines across nine genomic implementation research studies.Am J Hum Genet. 2020; 107: 932-941Abstract Full Text Full Text PDF PubMed Scopus (40) Google Scholar Patients, their families, and health care providers depend on precise classification of germline genetic variants to make possibly life-changing medical management decisions. Alongside the American College of Medical Genetics and Genomics and the Association for Molecular Pathology guidelines for variant interpretation,3Richards S. Aziz N. Bale S. Bick D. Das S. Gastier-Foster J. Grody W.W. Hegde M. Lyon E. Spector E. Voelkerding K. Rehm H.L. ACMG Laboratory Quality Assurance CommitteeStandards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology.Genet Med. 2015; 17: 405-424Abstract Full Text Full Text PDF PubMed Scopus (18634) Google Scholar a few public initiatives, such as the ClinVar database, have been launched to advance consistency and accuracy in germline variant classification. ClinVar is a free public archive of human genomic variants and interpretations of their relationships to human phenotypes. This database presents a main platform for objective evaluations of the interlaboratory reproducibility of variant classification and a procedure for discrepancy detection and resolution.4Landrum M.J. Chitipiralla S. Brown G.R. Chen C. Gu B. Hart J. Hoffman D. Jang W. Kaur K. Liu C. Lyoshin V. Maddipatla Z. Maiti R. Mitchell J. O'Leary N. Riley G.R. Shi W. Zhou G. Schneider V. Maglott D. Holmes J.B. Kattman B.L. ClinVar: improvements to accessing data.Nucleic Acids Res. 2020; 48: D835-D844Crossref PubMed Scopus (387) Google Scholar,5Frone M.N. Stewart D.R. Savage S.A. Khincha P.P. Quantification of discordant variant interpretations in a large family-based study of Li-Fraumeni syndrome.JCO Precis Oncol. 2021; 5PO.21.00320PubMed Google Scholar A variety of different submitters, including clinical testing laboratories, research groups, biocurators, expert panels, and practice guideline groups, have submitted their interpretations of variants, mostly germline, to the ClinVar database. Considering the large amount of data from many types of contributors in this database, it can be challenging for users to verify the validity of variant interpretations, especially if there are conflicting claims about the same variant. The ClinVar group does not modify interpretations, but rather combines submissions of the same variant and identifies whether interpretations are concordant.4Landrum M.J. Chitipiralla S. Brown G.R. Chen C. Gu B. Hart J. Hoffman D. Jang W. Kaur K. Liu C. Lyoshin V. Maddipatla Z. Maiti R. Mitchell J. O'Leary N. Riley G.R. Shi W. Zhou G. Schneider V. Maglott D. Holmes J.B. Kattman B.L. ClinVar: improvements to accessing data.Nucleic Acids Res. 2020; 48: D835-D844Crossref PubMed Scopus (387) Google Scholar,6Henrie A. Hemphill S.E. Ruiz-Schultz N. Cushman B. DiStefano M.T. Azzariti D. Harrison S.M. Rehm H.L. Eilbeck K. ClinVar miner: demonstrating utility of a Web-based tool for viewing and filtering ClinVar data.Hum Mutat. 2018; 39: 1051-1060Crossref PubMed Scopus (60) Google Scholar have ClinVar and variant to discordance in variant S. Lincoln S.E. Kobayashi Y. Nykamp K. Nussbaum R.L. Topper S. Sources of discordance among germ-line variant classifications in ClinVar.Genet Med. 2017; 19: 1118-1126Abstract Full Text Full Text PDF PubMed Scopus (75) Google L.M. J.J. Bale S. Hegde M. of genomic sequence to interpretation for J Hum Genet. Full Text Full Text PDF PubMed Scopus Google J. L. P. V. K. S.M. interpretation of genetic variants and commercial laboratories as the of Oncol. PubMed Scopus Google Scholar to different and there is the of S. Lincoln S.E. Kobayashi Y. Nykamp K. Nussbaum R.L. Topper S. Sources of discordance among germ-line variant classifications in ClinVar.Genet Med. 2017; 19: 1118-1126Abstract Full Text Full Text PDF PubMed Scopus (75) Google Scholar in showed that in submission and variant are to variant classification However, the of variants or more ClinVar the of there were such in the database In over of variants submitted to ClinVar are a single submitter are it would be for laboratories and to have an of the of variant classification. can be used to the accuracy or of variant classification This study a between variants with classification and with discordant classification to identify variant that be associated with classification discordance in ClinVar. this molecular as allele frequency, variant type, and variant and submission as the first submitter identity and were Furthermore, a set of using the Extreme Gradient Boosting to the of record classifications was may future research on classification concordance and may database users to using public The ClinVar and submission were from the ClinVar The E. P. and for the variant 2021; Scopus Google Scholar was used to the to a this variant such as variant variant ClinVar and in the and Sequencing the and the were In are if the discrepancy an interpretation of to variant of to or to a data set of multisubmitter variants ClinVar or ClinVar variants with the and were for the data set and a of variants were including and the data set was into and to the variant between The variants to ClinVar between and were used as the test data the study on variants record in that were to variants ClinVar ClinVar this this variants were into variants and to the database. and variants and conflict were in the test data make on the to submitter were from the ClinVar and as the and a of variants were in this variant effect Molecular W. L. S.E. G.R. A. P. The variant effect 17: PubMed Scopus Google Scholar was used to more to the data the data and the the were to consequence variant. The a of different including genomic and variant as as allele from and prediction are with prediction the for C. C. Y. Y. a database of and for human and Med. 2020; PubMed Scopus Google Scholar In are including and The predicted a of variant effect Medical and and J. P. M. A. K. Y. variant prediction with models of 2021; PubMed Scopus Google Scholar was to the data The of variant effect is a for the prediction of the of human variants that a on of J. P. M. A. K. Y. variant prediction with models of 2021; PubMed Scopus Google Scholar between the of variants, or was or was used to data for The of between was the or V. was as a of effect for the between The between was The was used to for the is to as and data was using the classification algorithm was used to make models to variants into conflict or The is as a and for in for C. a of the on and Association for Scopus Google Scholar Boosting is an in models are to the In such an models are in a improvements can be In the was with the of and C. a of the on and Association for Scopus Google Scholar The for using algorithm was the was large and there were of in the data The approach was to large data and the Chen and of C. a of the on and Association for Scopus Google Scholar is to on a data set with that the data from a of was allele the between different was In case of was for the of the Furthermore, the of between was into to to of the were as a of the of the is that it is A can be in different that the of the that the a of the and that the to be However, the of variants in the data set is of the in ClinVar from the of the data set can be with variants and conflict variants. for this the of algorithm was as in the J. D. D. a of in of for D. M. of the on of Scholar to train a of for for a of the detection for group with conflict was using the This the of group with to on allele from different including and the were in the ClinVar The data were from The between the conflict and in of of variant A was in the allele between the conflict and in the database, with the conflict group a allele with the group test The same was for the with a allele in the conflict group with the group test In the and showed the However, of the from and this data from and was to variants with allele in in the data set variants in and not in that variants different characteristics than in The ClinVar database genes harboring variants ClinVar and ClinVar The of variants gene with a of and a of variants. than genes variant. The and for the of variants in the genes were and to respectively. The genes with the of variants were and showed a between the of variant submissions for a gene and the of conflict interpretations in that as a of This that genes with a of variant submissions to have a of conflicting for the in the human the of variants in the data the genes the of conflict variants, with the of conflict variants in the data The types of that to classification discordance for variants in genes was A of conflict variants were in The of conflicting interpretation was to for of the to of the were between and for such as and the to was with more than of classification discordance to types of of in the of conflict this likely and are to be the as are likely and variant of in a In this likely and are to be the as are likely and P. variant of of the variants in the data set were variants. types of variants were in the data set and the between different variant variants, and variants the of and variants the of conflicting The data set different on the effect of the variants on the gene and The class of variants were and the were variants in the in the and the of conflicting In and variants the of conflicting was used to the a classification of the of the variant on with The in the in of are and The group was for the conflict the group was for conflict than different as the have been to the variants in the data of are and were to be to variants were between the in the conflict and of the variants with a were in the conflict of the variants that to were in the conflict group whether the of with variant classification the of variants was in different of the The is an approach to assess of the models for a There are in from of to of and for genes that not have in on the can be on the of the variants were or and was in the of variants into discordance However, for were to of variants were in the The study the between the variant class and of predicted a of prediction in C. C. Y. Y. a database of and for human and Med. 2020; PubMed Scopus Google Scholar and of variant J. P. M. A. K. Y. variant prediction with models of 2021; PubMed Scopus Google Scholar with of such as and as as that showed with algorithm were in the of the effect for such as and were among the as in of showed that the of conflict in the group was as as in the group A for was the of conflict variants in the group was more than as large as in the group The of the of different from that there was pronounced in the test effect the conflict group the group not from in a of variants the ClinVar database. were in the among of variants with Genetics as the first submitter were as a clinical across In from certain and There are of data in including clinical and among There was a in the of conflict variants between with research data associated with a of conflict with clinical testing were to models with in whether a variant would more likely be in the conflict or a was and the of the was This classifier an accuracy of on the set and on the The the was with a classification of the of algorithm was to the the optimized to the on the set and it in variants on the test the optimized a of with an accuracy of Furthermore, the weighted average of precision, recall, and f1-score were 0.84, 0.88, and 0.85, on the of variant conflicting of the and of variants in of the and of variants in average of the average of the to the of variants in Extreme Gradient the or is not The variant conflicting of the and of variants in average of the average of the to the of variants in in a Extreme Gradient the or is not The classifier variant class the the in variants were the first submitter identity, variant allele frequency, molecular submitter and models were to on the set from ClinVar. In a to a that is on labeled including and test The first was an classifier to the optimized but it was on data with class The and were to increase the detection for conflict variants. In the of models was for the conflict group the of the precision, in an increase in the The prediction of the models on variants can be provide into the were for the models on the train and test to the of the to in a the of variant advancements in public database variant effect prediction and the of conflicting clinical interpretation of variants a The of the and of conflicting with the to medical is crucial from a clinical J. L. P. V. K. S.M. interpretation of genetic variants and commercial laboratories as the of Oncol. PubMed Scopus Google Scholar The study to between variants that into conflict different characteristics of the variants in ClinVar. were to into the clinical classification discordance be predicted variant classifier models were used to the conflict class for variants in ClinVar. may provide a of in the of variants. Population allele conflict and variants in certain such as and a of and variants that with variants in data to have allele is as variants with allele are as between However, a allele was for conflict variants in and but not in that of variants in and were as variants in and were as and a of variants in This of variants in the and may the that not data in the and are from S. M. W. M. Rehm H.L. A. interpretation using from Mutat. 2021; PubMed Scopus Google B. M. S. of variants in the a of the of variant Mutat. Google Scholar variant it is to be of such genes have more submissions to ClinVar than This can be to a variety of such as the with genetic or the of certain in the of and research and M.J. J.M. Riley G.R. Jang W. Maglott D.R. ClinVar: public archive of relationships among sequence and human Acids Res. PubMed Scopus Google Scholar The of submissions of variants in gene was with the of conflict interpretations, with a of in genes more data from ClinVar used in this The of among with genes a than was not to that a large for the in the human L. The in and of and their to PubMed Scopus (75) Google Scholar a of conflicting variants, likely genes have more variants However, a the of conflict in the database, that the of conflict variants is not gene in are with a of in of Scholar There is for variants in and genes with a of conflict variants, such as and A. M. Molecular of J 2021; PubMed Scopus Google Scholar However, variants in genes that in of the of conflict are associated with different or of such genes are and among This that genetic and of be in the variant classification of the genes with the of conflict variants that of the were to variants that the to such types of conflict can be that for patients and there are for using and for medical management to this J. L. P. V. K. S.M. interpretation of genetic variants and commercial laboratories as the of Oncol. PubMed Scopus Google S.E. D. M. N. variant classification and for the interpretation of genetic test Mutat. PubMed Scopus Google Scholar in the conflict the to or to are of and showed the of types of genes are in and in can in There is a of variants can and modify the of in M. C. C. H. genotype is not of an of the molecular of in human Genet. PubMed Scopus Google M.J. Genetic of in Genet. 2020; PubMed Scopus Google Scholar genes are to a in in genes in Full Text Full Text PDF PubMed Scopus Google M. the of in the of PubMed Scopus Google Scholar This can lead to among with the same genotype in make it to the of a variant. to of conflict variants in were as or genetic may to a of conflicting of can in patients with variants This be the case for the N. P. H. M. L. P. J. B. J.J. gene PubMed Scopus Google Scholar of conflicting were between and of are to are gene of the variant were to of conflict on classifications of variants that in the of a of conflicting variant were in the with the This be to the that the of genes are more than the to more among in the classification of variants in the M. V. Genetic variants in 2018; PubMed Scopus Google N. S. M.J. S. J. S.A. the of variants in 2020; PubMed Scopus Google Scholar is to be the in human D. C. G.M. The of human genetic PubMed Scopus Google S. M. of in PubMed Scopus Google Scholar was in as the and were in the in the data and were associated with of conflict was that such as for and for were associated with of conflict This be a of the that variants are the American College of Medical Genetics and Genomics are large that a of data to ClinVar. variants have the to including and is that may lead to variants, to the of conflicting interpretations of variants. However, this study not a of and of variants to S. Lincoln S.E. Kobayashi Y. Nykamp K. Nussbaum R.L. Topper S. Sources of discordance among germ-line variant classifications in ClinVar.Genet Med. 2017; 19: 1118-1126Abstract Full Text Full Text PDF PubMed Scopus (75) Google Scholar showed that submissions for many of the classification The current that research submissions of the data conflict were in of the from this of the the of the first submitter identity was between clinical This is likely a of the interpretation of large variants to the database, the of the that group a on the may an to laboratories to provide the guidelines variant interpretation is to users in S. Aziz N. Bale S. Bick D. Das S. Gastier-Foster J. Grody W.W. Hegde M. Lyon E. Spector E. Voelkerding K. Rehm H.L. ACMG Laboratory Quality Assurance CommitteeStandards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology.Genet Med. 2015; 17: 405-424Abstract Full Text Full Text PDF PubMed Scopus (18634) Google K. M. M. J. B. Kobayashi Y. N. J. M. Topper S. Genomics a of the ACMG-AMP variant classification Med. 2017; 19: Full Text Full Text PDF PubMed Scopus Google Scholar data a of different features, including allele frequency, deleteriousness from different prediction and data on the first submitters, among were used to to a record into conflict or The data set used for this showed different for of the In the data are a the and implementation of models are not of working with data S. J. K. data is and in prediction using a 2021; Full Text Full Text PDF PubMed Scopus Google D. B. A on data in 2021; PubMed Scopus Google Scholar have been to with in the of an but are not and In this the of was an The first would a of and the would lead to S. J. K. data is and in prediction using a 2021; Full Text Full Text PDF PubMed Scopus Google Scholar data set was an especially and In it would be a and to of the models on such a data set Considering the models were using the is a of and is to models on data using an to C. a of the on and Association for Scopus Google Scholar in the and testing were using different of variants. is to that ClinVar ClinVar that were submitted in ClinVar. This approach was used to assess the of models on and data that variants. The and optimized on the However, the optimized classifier accuracy in variants on the test Considering the that genetic allele and variant on the and the first submitter can be used to the future conflict status of a single submitter variant. with different and were to provide a prediction of classification conflict for in ClinVar. variants predicted as especially ClinVar in the was was variants predicted as conflicting the models ClinVar in the was was was to of conflict but the of is to that the in this study are on and the classification of variants, including from single submitters, clinical and provide a more of the accuracy and of the There are to the current conflict variants were the conflict in more and conflict variants that consensus with or more S. Lincoln S.E. Kobayashi Y. Nykamp K. Nussbaum R.L. Topper S. Sources of discordance among germ-line variant classifications in ClinVar.Genet Med. 2017; 19: 1118-1126Abstract Full Text Full Text PDF PubMed Scopus (75) Google S.M. Chen W. Das S. J. H. R. R. K. N. S. Topper S. L.M. K. Rehm H.L. Variant of variant classification in ClinVar between clinical laboratories an Mutat. 2018; 39: PubMed Scopus Google Scholar the of the to variant classification such as the first submitter identity and the submitter was for models to make the first submitter identity was to the models for conflict class the of was set to this approach was used to provide for variants, the prediction may be to the submitter count to In the of the data set and for not be Furthermore, the of the ClinVar database and the of submissions over of the models would be to their accuracy and in the In this study showed that there are pronounced associations between variant classification discordance and variant allele frequency, the gene in the variant is and the identity of the first submitter to the ClinVar database. the predicted conflict class for single submitter variants deposited in the ClinVar database, may help variant or laboratories with their variant The from the current study may future research on classification concordance and may database users to using public with with with with used in the current used multisubmitter variants to train the to the and to an of the used single-submitter to provide the predicted conflict class for variants as a of classification database. with multisubmitter variants in ClinVar. Variant count for variants that have been to variants. The variants are the of conflicting variants from to The are a with the variant in in the the of variants in with between the variant class and prediction of different prediction the for an is on the The effect of and for test of is The the of in the data the of used in this to the M. and used the and used the and used the and used the and with of conflict and variants in different ClinVar. the the are the of submissions in that with of the Extreme Gradient Boosting classifier and the optimized with in the optimized
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".