Omics BioAnalytics: an RShiny application for multimodal biomarker panel discovery and assessment
Bibliographic record
Abstract
Motivation: Machine learning offers a powerful approach for building predictive models from high-dimensional molecular data. Omics technologies such as transcriptomics, proteomics, and metabolomics quantify thousands of molecules simultaneously, providing deep insights into disease biology. Integrating multiple modalities can enhance predictive performance, as shown in histology-omics and holter-omics applications. To support streamlined, reproducible, and user-friendly multimodal analytics, we developed Omics BioAnalytics, an R Shiny platform for unified analysis, integration, and interpretation of diverse omics datasets. Results: Omics BioAnalytics performs late integration using ensembles of elastic net models trained independently on each modality, with predictions averaged across datasets. The platform provides interactive dashboards for metadata exploration, exploratory analyses, differential expression, gene set analysis, and biomarker discovery. Results are visualized through dynamic plots and downloadable reports, ensuring transparent and reproducible workflows. A unique feature is the integrated multimodal Alexa Skill, which enables voice-based querying and rapid visualization. Together, these web and voice-enabled tools offer accessible and reproducible multimodal analytics for biomedical researchers, supporting the discovery of molecular signatures, predictive biomarkers, and therapeutic targets. Availability and implementation: All source code, public datasets, video walkthroughs, and the deployed application are available at: https://github.com/CompBio-Lab/omicsBioAnalytics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.046 | 0.025 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".