Validation of tissue microarray immunohistochemistry staining and interpretation in diffuse large B-cell lymphoma
Bibliographic record
Abstract
Tissue microarrays (TMAs) show concordance with whole tissue sections in the immunohistochemical evaluation of tumor cells. However, potential inter-institutional variability among observers and immunohistochemical staining methods has not been fully addressed. We selected 21 cases of diffuse large B-cell lymphoma (DLBCL) to process for TMAs. Immunohistochemical stains were performed in 3 laboratories, and reviewed independently by 3 hematopathologists at the 3 institutions. Stains were scored on a 4-point scale. Statistical analyses of variation in the scoring among observers and among different institutions' stains were performed. Stains for CD3, CD10, CD20, BCL-2, BCL-6, MIB-1, and FOX-P1 revealed little variation among observers, with an average 51-82% complete agreement and 82-100% agreement +/- 1 numerical score. The rate of concordance when evaluating most stains performed in different laboratories was also relatively good, with an average of 55-72% complete agreement and 70-97% agreement +/- 1 score. However, scoring of MUM-1 and p53 stains showed wider variation, with an average of only 37 and 30% complete agreement among observers, and 11 and 45% agreement when stains from different institutions were examined. Further statistical analyses were performed to compare the observers' scoring of their own institution's stains (self-review) vs. observers' scoring of other institutions' stains (non-self). The agreement rate for the p53 stain was significantly higher when based on self-review (average 58% complete agreement) compared with an agreement rate of only 10.5% when based on a review of stains performed in another laboratory, non-self review, P < 0.01. This difference in the self- vs. non-self review was not seen when data for MUM-1 were analysed. In conclusion, most phenotypic markers used in the analysis of DLBCL can be evaluated in TMAs with adequate agreement among observers and laboratories. These include CD3, CD20, CD10, BCL-2, BCL-6, MIB-1, and FOX-P1. However, some markers, such as p53 and MUM-1, are more prone to inter-institutional variation. Variations in interpretation can be partially overcome by self-adjusted/adapt tendency, as seen with p53. Especially with newly developed markers, such as MUM-1, the development of standardized techniques for staining and interpretation is critical to reduce inter-observer variability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".