OvMark: a user-friendly system for the identification of prognostic biomarkers in publically available ovarian cancer gene expression datasets
Bibliographic record
Abstract
BACKGROUND: Ovarian cancer has the lowest survival rate of all gynaecologic cancers and is characterised by a lack of early symptoms and frequent late stage diagnosis. There is a paucity of robust molecular markers that are independent of and complementary to clinical parameters such as disease stage and tumour grade. METHODS: We have developed a user-friendly, web-based system to evaluate the association of genes/miRNAs with outcome in ovarian cancer. The OvMark algorithm combines data from multiple microarray platforms (including probesets targeting miRNAs) and correlates them with clinical parameters (e.g. tumour grade, stage) and outcomes (disease free survival (DFS), overall survival). In total, OvMark combines 14 datasets from 7 different array platforms measuring the expression of ~17,000 genes and 341 miRNAs across 2,129 ovarian cancer samples. RESULTS: To demonstrate the utility of the system we confirmed the prognostic ability of 14 genes and 2 miRNAs known to play a role in ovarian cancer. Of these genes, CXCL12 was the most significant predictor of DFS (HR = 1.42, p-value = 2.42x10-6). Surprisingly, those genes found to have the greatest correlation with outcome have not been heavily studied in ovarian cancer, or in some cases in any cancer. For instance, the three genes with the greatest association with survival are SNAI3, VWA3A and DNAH12. CONCLUSIONS/IMPACT: OvMark is a powerful tool for examining putative gene/miRNA prognostic biomarkers in ovarian cancer (available at http://glados.ucd.ie/OvMark/index.html). The impact of this tool will be in the preliminary assessment of putative biomarkers in ovarian cancer, particularly for research groups with limited bioinformatics facilities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.012 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.024 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".