bayesReact: Expression-coupled regulatory motif analysis detects microRNA activity in cancer and at the single cell level
Bibliographic record
Abstract
Motivation: Regulatory constraints are crucial in maintaining tissue and cell integrity, and play important roles during developmental processes and environmental responses. Yet many regulatory mechanisms remain unobserved at the single-cell level and statistical inference may, in some cases, help elucidate their condition-specific activity and perturbation during disease progression. Results: We introduce bayesReact (BAYESian modeling of Regular Expression ACTivity), a generative model of motif occurrence across experimentally ranked sequences to infer motif-based regulatory activities. The method is evaluated for microRNAs (miRNAs), which perform post-transcriptional regulation through target mRNA destabilization and translational repression. Inferred miRNA activities positively correlate with the observed miRNA expressions in primary tumors from The Cancer Genome Atlas (TCGA) and mouse stem cells. The top miRNA activity profiles are as informative for TCGA cancer-type cluster identification as the top miRNA or mRNA expression profiles. The activity captures tissue-specific miRNA patterns observed in the matched expression, e.g., the expression of miR-122-5p in the liver and miR-124-3p in low-grade gliomas (LGG). We observe a negative association between the activity of the two miRNAs and their target gene expressions, including between the miR-124-3p activity and the anti-neuronal REST expression in LGG. bayesReact outperforms the existing method, miReact, on sparse count data, and shows a higher correlation with the miRNA expression in single-cell data. The method recovers temporal activities of prominent miRNAs during murine stem cell differentiation, including miR-298-5p, miR-92-2-5p, and the large Sfmbt2 cluster (miR-297-669). The bayesReact model is probabilistic and quantifies the uncertainty of all provided estimates. It is unsupervised and permits screens of bulk or single-cell data to identify condition-specific regulatory motif candidates. It further improves miRNA activity inference in single-cell data. Availability and implementation: bayesReact is implemented as an R-package, uses a Hamiltonian Monte-Carlo sampler for posterior approximation, and is available at https://github.com/astamr/bayesReact.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".