Assessing Oral Epithelial Dysplasia Risk for Transformation to Cancer: Comparison Between Histologic Grading Systems Versus S100A7 Immunohistochemical Signature-based Grading
Bibliographic record
Abstract
While a 3-tier oral epithelial dysplasia grading system has been utilized for decades, it is widely recognized as a suboptimal risk indicator for transformation to cancer. A 2-tier grading system has been proposed, although not yet validated. In this study, the 3-tier and 2-tier dysplasia grading systems, and an S100A7 immunohistochemical signature-based grading system were compared to assess prediction of risk of transformation to oral cancer. Formalin-fixed, paraffin-embedded biopsy specimens with known clinical outcomes were obtained retrospectively from a cohort of 48 patients. Hematoxylin and eosin-stained slides were used for the 2- and 3-tier dysplasia grading, while S100A7 for biomarker signature-based assessment was based on immunohistochemistry. Inter-observer variability was determined using Cohen's kappa ( K ) statistic with Cox regression disease free survival analysis used to determine if any of the methods were a predictor of transformation to oral squamous cell carcinoma. Both the 2- and 3-tier dysplasia grading systems ranged from slight to substantial inter-observer agreement ( Kw between 0.093 to 0.624), with neither system a good predictor of transformation to cancer (at least P =0.231; ( P >>>0.05). In contrast, the S100A7 immunohistochemical signature-based grading system showed almost perfect inter-observer agreement ( Kw =0.892) and was a good indicator of transformation to cancer ( P =0.047 and 0.030). The inherent grading challenges with oral epithelial dysplasia grading systems and the lack of meaningful prediction of transformation to carcinoma highlights the significant need for a more objective, quantitative, and reproducible risk assessment tool such as the S100A7 immunohistochemical signature-based system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".