DRUGPATH: A New Database for Mapping Polypharmacology
Bibliographic record
Abstract
While there are existing databases that curate only drug, target, or pathway data for instance, none of these alone are exhaustive. The Drug Gene Pathway (DRUGPATH) meta database was created as a response to the complex treatment required for various diseases including Gulf War Illness (GWI) and post-traumatic stress disorder (PTSD), where therapy involves using multiple drugs in combination. Here, drug-drug interactions can occur due to the promiscuous nature of pharmaceuticals, which can then lead to various side effects or can alternatively be utilized towards drug repurposing. The objective was to develop a database that maps the interactions between drugs, genes, pathways, and targets for use in the treatment of complex diseases, including the prediction of off-target interactions, otherwise known as side effects. Using MATLAB and Python scripts, interactions between known drugs, genes, targets, and pathways amalgamated from numerous expert-curated sources such as PharmGKB, DrugBank, DGIdb, ConsesusPathDB, Guide to PHARMACOLOGY, HUGO Gene Nomenclature Committee, Toxin and Toxin-Target Database, repoDB, the FDA’s National Drug Code database, etc. were mapped together. The raw data was first downloaded from its source and subsequently cleaned, where extraneous information such as data from non-humans, internal identifiers, timestamps, etc. were removed. The remaining information was then integrated into an SQLite database. DRUGPATH currently contains a total of 2,632,516 unique entries, and of these, there are 54,757 unique genes, 2,632,242 unique pathways, and 31,042 unique drugs. DRUGPATH allows researchers and clinicians to discern which pathways are affected by each drug, reducing the likelihood of an adverse drug reaction occurring. The incorporation of drug, gene, target, and pathway information makes DRUGPATH a powerful resource for predicting potential side effects when designing or refining a given drug combination therapy. Not only that, but we have additionally added the FDA status, half-life, and indication for each drug whenever possible for clinical applications of this database.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".