The reliability of immunohistochemical analysis of the tumor microenvironment in follicular lymphoma: a validation study from the Lunenburg Lymphoma Biomarker Consortium
Bibliographic record
Abstract
The cellular microenvironment in follicular lymphoma is of biological and clinical importance. Studies on the clinical significance of non-malignant cell populations have generated conflicting results, which may partly be influenced by poor reproducibility in immunohistochemical marker quantification. In this study, the reproducibility of manual scoring and automated microscopy based on a tissue microarray of 25 follicular lymphomas as compared to flow cytometry is evaluated. The agreement between manual scoring and flow cytometry was moderate for CD3, low for CD4, and moderate to high for CD8, with some laboratories scoring closer to the flow cytometry results. Agreement in manual quantification across the 7 laboratories was low to moderate for CD3, CD4, CD8 and FOXP3 frequencies, moderate for CD21, low for MIB1 and CD68, and high for CD10. Manual scoring of the architectural distribution resulted in moderate agreement for CD3, CD4 and CD8, and low agreement for FOXP3 and CD68. Comparing manual scoring to automated microscopy demonstrated that manual scoring increased the variability in the low and high frequency interval with some laboratories showing a better agreement with automated scores. Manual scoring reliably identified rare architectural patterns of T-cell infiltrates. Automated microscopy analyses for T-cell markers by two different instruments were highly reproducible and provided acceptable agreement with flow cytometry. These validation results provide explanations for the heterogeneous findings on the prognostic value of the microenvironment in follicular lymphoma. We recommend a more objective measurement, such as computer-assisted scoring, in future studies of the prognostic impact of microenvironment in follicular lymphoma patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.038 | 0.038 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".