Separating the Wheat from the Chaff: The Use of Upstream Regulator Analysis to Identify True Differential Expression of Single Genes within Transcriptomic Datasets
Bibliographic record
Abstract
The development of DNA microarray and RNA-sequencing technology has led to an explosion in the generation of transcriptomic differential expression data under a wide range of biologic systems including those recapitulating the monogenic muscular dystrophies. Data generation has increased exponentially due in large part to new platforms, improved cost-effectiveness, and processing speed. However, reproducibility and thus reliability of data remain a central issue, particularly when resource constraints limit experiments to single replicates. This was observed firsthand in a recent rare disease drug repurposing project involving RNA-seq-based transcriptomic profiling of primary cerebrocortical cultures incubated with clinic-ready blood–brain penetrant drugs. Given the low validation rates obtained for single differential expression genes, alternative approaches to identify with greater confidence genes that were truly differentially expressed in our dataset were explored. Here we outline a method for differential expression data analysis in the context of drug repurposing for rare diseases that incorporates the statistical rigour of the multigene analysis to bring greater predictive power in assessing individual gene modulation. Ingenuity Pathway Analysis upstream regulator analysis was applied to the differentially expressed genes from the Care4Rare Neuron Drug Screen transcriptomic database to identify three distinct signaling networks each perturbed by a different drug and involving a central upstream modulating protein: levothyroxine (DIO3), hydroxyurea (FOXM1), dexamethasone (PPARD). Differential expression of upstream regulator network related genes was next assessed in in vitro and in vivo systems by qPCR, revealing 5× and 10× increases in validation rates, respectively, when compared with our previous experience with individual genes in the dataset not associated with a network. The Ingenuity Pathway Analysis based gene prioritization may increase the predictive value of drug–gene interactions, especially in the context of assessing single-gene modulation in single-replicate experiments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.005 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".