Histo- and immunohistochemistry-based estimation of the TCGA and ACRG molecular subtypes for gastric carcinoma and their prognostic significance: A single-institution study
Bibliographic record
Abstract
Gastric cancers comprise molecularly heterogeneous diseases; four molecular subtypes were identified in the cancer genome atlas (TCGA) study, with implications in patient management. In our efforts to devise a clinically feasible means of subtyping, we devised an algorithm based on histology and five stains available in most academic pathology laboratories. This algorithm was used to subtype our cohort of 107 gastric cancer patients from a single institution (St. Michael's Hospital, Toronto, Canada), which was divided into 3 cases of EBV-positive, 23 of MSI, 27 of GS and 54 of CIN tumours. 87% of the tumours with diffuse histology were classified as GS subtype, which was notable for younger age. Examining for characteristic molecular features, aberrant p53 immunostaining was seen most frequently in the CIN subtype (43% in CIN vs. 6% in others), whereas ARID1A loss was rarely seen (6% vs. 35% in others). HER2 overexpression was seen exclusively in CIN tumours (17% of CIN tumours). PD-L1 positivity was seen predominantly in the EBV and MSI tumours. As with the TCGA study, no survival differences were seen between the subtypes. A similar strategy was employed to approximate the Asian Cancer Research Group (ACRG) molecular subtyping, with the addition of p53 IHC to the algorithm. We observed rates of ARID1A loss and HER2 overexpression that were comparable to the ACRG study. In summary, our algorithm allowed for clinically feasible means of subtyping gastric carcinoma that recapitulated the key molecular features reported in the large scale studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".