Cross-Platform Methylation-Based Site of Origin Classification for Squamous Cell Carcinomas
Bibliographic record
Abstract
Squamous cell carcinomas (SCCs) are one of the most common cancer types and can arise at nearly any anatomic site. Because SCCs are one of the most common metastases, do not have reliable site-specific morphologic or genomic features, and have considerable morphologic and immunohistochemical overlap with urothelial carcinomas, distinguishing between primary and metastatic squamous-appearing tumors can be challenging. This distinction can be critical to clinical management. We present Squamous cell carcinoma Methylation for Origin Site (SquaMOS), a methylation-based classifier to predict site of origin of squamous-appearing carcinomas. Trained on publicly available array-based methylation data from 1062 primary SCCs (from lung, head and neck, cervix, and esophagus) and urothelial carcinomas, SquaMOS predicted site of origin in primary tumors with 96.1% accuracy in an internal test set (n = 458) and 97.4% accuracy in an external test set from 3 institutions (n = 78). On metastatic tumors (n = 51), SquaMOS predictions were 96.1% accurate. SquaMOS was directly applicable to shallow Nanopore sequencing data (CpG probe site coverage, 0.25-2.88×) with an accuracy of 91.7% (n = 36; 100% accurate for high-confidence predictions). When tested on SCCs outside the training set types (n = 15, including 3 metastases to lung), no cases were misclassified as of lung origin, supporting accuracy of lung vs nonlung origin classification for diverse SCC types. Overall, we demonstrate highly accurate performance of the SquaMOS classifier on primary and metastatic tumors from multiple data sources, robust to suboptimal tumor purity. We illustrate transferability of our array-based classifier to low-depth Nanopore sequencing data, a potentially rapid means of site of origin determination in a clinical setting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".