Prediction of gene-based drug indications using compendia of public gene expression data and PubMed abstracts
Bibliographic record
Abstract
The tremendous research effort on diseases and drug discovery has produced a huge amount of important biomedical information which is mostly hidden in the web. In addition, many databases have been created for the purpose of storing enormous amounts of information and high-throughput experiments related to drugs and diseases' effects on genes. Thus, developing an algorithm to integrate biological data from different sources forms one of the greatest challenges in the field of computational biology. Based on our belief that data integration would result in better understanding for the drug mode of action or the disease pathophysiology, we have developed a novel paradigm to integrate data from three major sources in order to predict novel therapeutic drug indications. Microarray data, biomedical text mining data, and gene interaction data have been all integrated to predict ranked lists of genes based on their relevance to a particular drug or disease molecular action. These ranked lists of genes have finally been used as a raw material for building a disease-drug connectivity map based on the enrichment between the up/down tags of a particular disease signature and the ranked lists of drugs. Using this paradigm, we have reported 13% sensitivity improvement in comparison with using microarray or text mining data independently. In addition, our paradigm is able to predict many clinically validated disease-drug associations that could not be captured using microarray or text mining data independently.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".