Abstract 94: DrugPath: A web database for academic investigators to identify drugs that match oncology molecular targets
Bibliographic record
Abstract
Abstract High-throughput genomics can identify oncology molecular targets, but translation into novel therapeutics requires access to relevant drugs. Pharmaceutical pipeline databases, such as MedTRACK (www.medtrack.com), can identify such drugs, but these databases are geared for commercial operations and present significant cost barriers to non-profit academic research. We have developed the DrugPath database (www.drugpath.org) as a comprehensive, free-of-charge resource for academic investigators. DrugPath can identify drugs within the cancer pipeline that may act against molecular targets of interest. DrugPath searchable data include names and mechanisms of drugs in the oncology pipeline (organized by clinical trial phase and route of administration), drug sponsor, and also tested combinations with other drugs. Drugs and drug sponsors are initially identified from ClinicalTrials.gov, biotech company websites, and other online resources; specific pipeline information is extracted by DrugPath personnel according to written protocols. PERL scripts to automate a portion of the data mining are being developed, although manual verification will remain to ensure data accuracy and relevance. The website is authored in XHTML and PHP, and content is stored in a MySQL database. We compared the extent of the DrugPath database to that of MedTRACK. DrugPath covers more than 475 companies, or 50% of 938 companies with oncology therapeutics in development as identified by MedTRACK. A search of DrugPath for inhibitors of VEGF, a well-established target, yielded 58 lead drug candidates (17% of 332 drugs identified by MedTRACK) from 45 total companies (47% of 95 companies identified by MedTRACK). A search for Hsp90 inhibitors yielded 15 companies (78% of 19 in MedTRACK) and 17 lead drug candidates (30% of 56 in MedTRACK), and a search for aurora kinase inhibitors yielded 17 companies (94% of 18 in MedTRACK) and 17 lead drug candidates (70% of 24 in MedTRACK). Ongoing efforts to populate the DrugPath database are focused on identifying drug development programs for emerging drug targets rather than expanding existing coverage of drug candidates against well-known targets. To support the end-goal of DrugPath–to facilitate obtaining drugs for academic research–DrugPath contains contact information (when available) for each company to assist with the process of developing material transfer agreements (MTAs). This database should be useful to academic investigators studying particular molecular targets and also to clinical investigators seeking to develop novel clinical trials. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 101st Annual Meeting of the American Association for Cancer Research; 2010 Apr 17-21; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2010;70(8 Suppl):Abstract nr 94.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.014 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.009 | 0.008 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.108 | 0.095 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".