Robust prognostic value of a knowledge-based proliferation signature across large patient microarray studies spanning different cancer types
Bibliographic record
Abstract
Tumour proliferation is one of the main biological phenotypes limiting cure in oncology. Extensive research is being performed to unravel the key players in this process. To exploit the potential of published gene expression data, creation of a signature for proliferation can provide valuable information on tumour status, prognosis and prediction. This will help individualizing treatment and should result in better tumour control, and more rapid and cost-effective research and development. From in vitro published microarray studies, two proliferation signatures were compiled. The prognostic value of these signatures was tested in five large clinical microarray data sets. More than 1000 patients with breast, renal or lung cancer were included. One of the signatures (110 genes) had significant prognostic value in all data sets. Stratifying patients in groups resulted in a clear difference in survival (P-values <0.05). Multivariate Cox-regression analyses showed that this signature added substantial value to the clinical factors used for prognosis. Further patient stratification was compared to patient stratification with several well-known published signatures. Contingency tables and Cramer's V statistics indicated that these primarily identify the same patients as the proliferation signature does. The proliferation signature is a strong prognostic factor, with the potential to be converted into a predictive test. Furthermore, evidence is provided that supports the idea that many published signatures track the same biological processes and that proliferation is one of them.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.021 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".