Pathology the Gold Standard – A Retrospective Analysis of Discordant “Second-Opinion” Lymphoma Pathology and Its Impact on Patient Care.
Bibliographic record
Abstract
Abstract A change in diagnosis on review of pathology, “second-opinion pathology”, is not uncommon for hematological malignancies with a range of between 11–20%. Hence, there is a significant potential for incurring a diagnostic error with implications during the clinical management of patients. We conducted a retrospective review to evaluate both potential and actual medical error noted in the management of lymphoma patients treated at Princess Margaret Hospital (PMH), by evaluating results of pathological review and identifying patients with a change in diagnosis between the initial referring centre and PMH (i.e. discordant pathology). All consecutive cases seen in medical or radiation oncology clinics between January 01, 2000 to June 30, 2003 with follow-up data collected to Dec 31, 2003 were evaluated. Only patients who had lymphoma first diagnosed after January 01, 2000 who had a referring centre pathology report and had clinical care at PMH were included. RESULTS: There were 2818 consecutive lymphoma patients identified of whom only 1065 (38%) met inclusion/exclusion criteria. There were 176 cases with discordant pathology identified in 171 individual patients (discordance rate of 16%); specimens evaluated were from nodal tissue – 129, extra-nodal – 36 and bone marrow – 11. The most common reasons for discordance were: malignant ↔ non-malignant – 27 cases, Non-Hodgkins ↔ Hodgkins – 14 cases, lymphoma ↔ solid tumour – 18 cases and more aggressive lymphoma ↔ less aggressive lymphoma – 47 cases. We found that disagreement in morphology was most often responsible for change (40%) followed by morphology and immunohistochemistry (27%). Cutaneous biopsies were found to have a higher rate of discordance than other biopsy sites. The 176 cases were graded by 6 blinded reviewers (pathologists and clinicians as well as physicians not affiliated with PMH) on a scoring system from minimal to severe with respect to potential for harm. Grading: not significant = 20 cases, minimal = 38 cases, moderate = 73 cases and severe = 43 cases. Overall, 66% of cases were deemed to have a moderate to severe potential for harm based on their discordant pathology. For these discordant cases, actual clinical management was based on PMH pathology interpretation in 52% versus 2% who were treated based referring centre diagnosis. 63 of 176 cases (37%) required additional biopsies or a more definitive biopsy to resolve the discordance. This resulted in 21 patients having significant surgical procedures such as partial gastrectomies, lumpectomies and mediastinal surgery. Overall, based on treatment policies and the treating physician’s opinion, there were 16 patients who were over-treated, 4 patients under-treated, 9 patients who had a significant change in planned treatment and 2 patients incorrectly treated. CONCLUSIONS: a discordance rate of 16% was similar to previous studies and this high rate maybe improved through centralization of lymphoma pathology;these types of patients are clearly at risk for harm, as best exemplified by patients who were felt to have a benign pathology that was actually malignant;Discordant pathology has clear clinical implications including serial biopsies, invasive testing and treatment delays.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.024 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".