Accuracy of kidney cancer diagnosis and histological subtype within Canadian cancer registry data
Bibliographic record
Abstract
INTRODUCTION: Provincial/territorial cancer registries (PTCRs) are the mainstay for Canadian population-based cancer statistics. Each jurisdiction captures this data in a population-based registry, including the Nova Scotia Cancer Registry (NSCR). The goal of this study was to describe data from the NSCR regarding renal cell carcinoma (RCC) pathology subtype and method of diagnosis and compare it to the actual pathology reports to determine the accuracy of diagnosis and histological subtype assignment. METHODS: This retrospective analysis included patients diagnosed with RCC in the NSCR from 2006-2010 with an ICD-O-3 code C64.9 seen or treated in the largest NS health district. From the NSCR, method of diagnosis and pathological diagnosis was recorded. All diagnoses of non-clear-cell RCC (nonccRCC) from NSCR were compared to the actual pathology report for descriptive comparison and reasons for discordance. RESULTS: 723 patients make up the study cohort. 81.3% of patients were diagnosed by nephrectomy, 11.1% radiography, 6.9 % biopsy, and 0.7% autopsy. By NSCR data, 52.8% had clear-cell (ccRCC), 20.5% RCC not otherwise specified (NOS), 12.7% papillary, 4% chromophobe, and the rest had other nonccRCC subtypes. By pathology reports, 69.5% had clear-cell, 15% papillary, 5% chromophobe, only 2.7% RCC NOS. There was a discordance rate of 15.4% between NSCR data and diagnosis from pathology report. Reasons for discordance were not enough information by the pathologist in 45.5%, misinterpretation of report by data coder in 22.2%, and true coding error in 32.3%. CONCLUSIONS: When using PTCR for RCC incidence data, it is important to understand how the diagnosis is made, as not all are based on pathological confirmation; in this cohort 11% were based on radiology. One must also be aware that clear-cell and non-clear-cell subtypes may differ between the PTCR data and pathology reports. In this study, ccRCC made up 52.8% of the registry diagnoses, but increased to 69.6% on pathology report review. Use of synoptic reporting and ongoing education may improve accuracy of registry data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.037 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".