Effect of automated TIL quantification in early-stage melanoma on accuracy of standard T staging using AJCC guidelines.
Bibliographic record
Abstract
10076 Background: Patients diagnosed with early stage melanoma are at risk of recurrence and death. Adjuvant therapy decreases risk but incurs toxicity and expense. While tumor-infiltrating lymphocytes (TILs) improve prognosis, studies have shown conflicting results due, at least in part, to inter-observer variability. Thus, TILs are not included in standard American Joint Committee on Cancer (AJCC) staging. Here, we quantitatively analyze TILs in hematoxylin and eosin (H&E) melanoma images using two machine learning algorithms. Methods: H&E images were evaluated by two methods for patients with resectable stage I-III melanoma from Columbia (N = 81) and validated using samples from Geisinger and Moffitt (N = 128). For both methods, H&E images were manually annotated using open source software, QuPath, to specify tumor regions. For Method A, images were divided into patches and, for each patch, a probability was generated to detect lymphocytes. Patches above a set threshold were considered to be “TIL positive”. Ratio of TIL positive patches to total patches was assessed for every image. For Method B, a classifier was manually trained in QuPath and then applied on each image to determine the ratio of the areas of all immune cells to all tumor cells as previously published. Cutoff values to define high and low risk groups were established based on a test set and then validated in an independent cohort. Results: Both methods distinguished patients with visceral recurrence from those without for the Columbia training set (Method A p = .0015, Method B p = .043). Using Method A, Kaplan-Meier curve at the selected cutoff also correlated significantly with disease specific survival (DSS) for Columbia (p = .022) and was validated in the Geisinger/Moffitt (p = .046) cohort. Cox analysis using Method A showed that TIL status predicted DSS in the validation set (p = .047) and added significantly to depth and ulceration (HR = 3.43, CI: 1.047-11.257, p = .042). Conclusions: Both open source machine learning algorithms find significantly higher TILs in patients who do not develop metastasis. Notably, Method A may add to standard predictors, such as depth and ulceration. These results demonstrate the promise of computational algorithms to enhance visual grading, and suggest that digital TIL evaluation may add to current AJCC staging. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".