MétaCan
Menu
← Back to cohort
Record W3083298312 · doi:10.1158/1538-7445.am2020-3211

Abstract 3211: Evolution of the CIViC knowledgebase for community driven curation of clinical variants in cancer

2020· article· en· W3083298312 on OpenAlexaff
Arpad Danos, Kilannin Krysiak, Erica K. Barnell, Adam Coffman, Joshua F. McMichael, Susanna Kiwala, Nicholas C. Spies, Lana Sheta, Shahil Pema, Lynzey Kujan, Kaitlin A. Clark, S. Anderson, Amber Z. Wollam, Brian Li, Justin Guerra, Shruti Rao, Deborah Ritter, Cameron J. Grisdale, Gordana Raca, Alex H. Wagner, Subha Madhavan, Malachi Griffith, Obi L. Griffith

Bibliographic record

VenueCancer Research · 2020
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsBC Cancer Agency
Fundersnot available
KeywordsData curationAnnotationInteroperabilityWorld Wide WebComputer scienceDigital curationBottleneckData scienceArtificial intelligence

Abstract

fetched live from OpenAlex

Abstract With increasing adoption of next generation sequencing into clinical practice, the problem of clinical interpretation of tumor variants arises as a bottleneck for patient care, as effective annotation of these variants draws from a constantly increasing body of largely unstructured clinical and preclinical results. One model to sustain clinical variant curation is to house variant annotation behind a paywall, using access fees to fund further curation effort. An alternate approach is to leverage public curation and expert moderation to create a free public resource to house and distribute this knowledge. The Clinical Interpretation of Variants in Cancer (CIViC, www.civicdb.org) knowledgebase employs the latter approach, and is a free and open-access public resource with an intuitive user interface and flexible public API for programmatic access to all content, which is available to the public with no restrictions on usage. All data is available for retrieval without login, while registration with a free account is required to contribute curation. The provenance of all curation and revision in CIViC is viewable through the web interface, and curators may also leave public comments on all content. Selected expert editors review and revise submitted content, which is clearly labeled as accepted once it has fully undergone moderation. All content in CIViC adheres to a structured data model which follows a published standard operating procedure for curation. This data model incorporates ontologies, standards and guidelines from across the field to promote interoperability and compatibility with other efforts. The CIViC interface also allows curators and organizations to track and display summary statistics of all their activity. CIViC currently has a community of over 190 curators and 16,000 clinical and research users around the world. CIViC continually develops and improves both new and existing features in response to user feedback as well as collaborative and internal development goals. Recently, the drugs and treatment terms used in predictive/therapeutic annotation have been normalized to the NCI Thesaurus. A conflict of interest (COI) statement is now required for all CIViC Editors, and functionality for writing and displaying the COI has been built into the interface. CIViC employs Predictive (Therapeutic), Prognostic, Diagnostic, Predisposing evidence types, and we highlight the recently introduced Functional evidence type, which has seen continued development. We will present the rationale for these changes including demonstrating how adding a Dominant Negative term better supports curation of functional genomics data sets. A focus of functional curation has been TP53, with over 50 evidence items to date. With multiple use cases for this type of data including targeted therapeutics, identification of relevant hotspots, or characterization of cancer driver mechanisms, functional evidence can be used to support conventional concepts of clinical utility and expand the CIViC data model. These developments provide a mechanism for discussion and integration of functional data into somatic variant interpretation guidelines, an area being explored but lacking expert consensus. Citation Format: Arpad Danos, Kilannin Krysiak, Erica K. Barnell, Adam C. Coffman, Joshua F. McMichael, Susanna Kiwala, Nicholas C. Spies, Lana M. Sheta, Shahil P. Pema, Lynzey Kujan, Kaitlin A. Clark, Sydney Anderson, Amber Wollam, Brian Li, Justin Guerra, Shruti Rao, Deborah I. Ritter, Cameron J. Grisdale, Gordana Raca, Alex H. Wagner, Subha Madhavan, Malachi Griffith, Obi L. Griffith. Evolution of the CIViC knowledgebase for community driven curation of clinical variants in cancer [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 3211.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.029
metaresearch head score (Gemma)0.090
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.029
Threshold uncertainty score0.154

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0290.090
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0130.011
Science and technology studies0.0030.002
Scholarly communication0.0130.009
Open science0.0090.013
Research integrity0.0060.005
Insufficient payload (model declined to judge)0.0260.030

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.184
GPT teacher head0.476
Teacher spread0.291 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueCancer Research→Same topicCancer Genomics and Diagnostics→French-language works237,207→