Bibliographic record
Abstract
Cancer is a disease driven by aberrant gene expression leading to uncontrolled cellular proliferation. To date, tens-of-thousands of patient tumors have undergone RNA-sequencing, creating a catalog of the genes and pathways differentially expressed in cancer tissues. Despite these significant advances, a fundamental aspect of gene regulation remains poorly characterized: the overall expression level across all genes. Recent work has demonstrated that certain oncogenes, such as MYC, might drive tumor growth by globally increasing transcription of all active genes, a phenomenon known as hypertranscription. While hypertranscription has been studied in model systems and cell lines, where drugs that dampen global transcription have shown promise against aggressive ‘transcriptionally addicted’ cancers, hypertranscription has never been characterized in human patients. Thus, we do not know hypertranscription’s prevalence across cancer types, its drivers, or its impact on patient outcomes. This gap in knowledge is driven in large part by the absence of appropriate methods to accurately measure global transcription. Nearly all reported gene expression estimates incorrectly assume relatively equal RNA output across samples. In Chapter 2, a novel computational method is developed that allows joint measurement of global and focal gene expression changes in patient tumor RNA-sequencing data. Critically, this method accounts for differences in tumor purity and ploidy, providing a direct fold-change measure in overall cancer-cell transcription. In Chapter 3, this method is applied to 7,494 tumor samples spanning 31 types, revealing that hypertranscription is a hallmark of aggressive cancers, with over 40% of all cancers harboring hypertranscription levels of at least 2-fold. Investigation of single-cell RNA-sequencing data revealed hypertranscriptional clones that dominated transcript production regardless of their size. Exploration of transcription factors revealed that loss of transcriptional suppression may be fundamental to the hypertranscriptional phenotype. In Chapter 4, the clinical implications of hypertranscription are explored. Hypertranscription defined patient subgroups with worse survival across multiple cancers, even within well-established subtypes. Finally, patients with hypertranscribed mutations have improved response to immune checkpoint therapy. Taken together, this work provides fundamental insights into gene dysregulation across human cancers and may prove useful in identifying patients that would benefit from novel therapies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".