Accuracy of kidney cancer diagnosis and histological subtype within cancer registry data.
Bibliographic record
Abstract
606 Background: Cancer registries are the mainstay for population-based cancer statistics including incidence and cancer type. In Canada, each province captures this data in provincial registries including the Nova Scotia Cancer Registry (NSCR).The goal of this study was to describe data from the NSCR about method of diagnosis and kidney cancer (KC) pathology and compare it to the actual pathology reports to determine the accuracy of diagnosis and histological subtype assignment. Methods: This retrospective analysis included patients with KC in the NSCR with an ICD-10-CM code C64.9 (malignant neoplasm of unspecified kidney, except renal pelvis) within the largest provincial metropolitan area from 2006-2010. Method of KC diagnosis (clinical, radiologic, histology, or autopsy) was recorded as was the pathological diagnosis based on WHO classification. All non-clear cell KC (nonccKC) diagnosis from the registry were compared to the actual pathology report (and pathology re-review when necessary) for comparison. Results: 733 pts make up the study cohort. 81.2% of patients were diagnosed based on nephrectomy, 11.5% on radiography, 6.5 % biopsy, and 0.8% autopsy. By registry data 53.1% had clear cell, 20.2% KC not otherwise specified (NOS), 12.7% papillary, 3.8% chromophobe, and many other nonccKC. By pathology reports, 62.2% had clear cell, 13.4% papillary, 4.4% chromophobe, only 2% KC NOS (because most radiological diagnosis were classified this way). A large number of pathological diagnoses make up the other nonccKC and discrepancies between registry data and pathology reports will be described and compared in detail. Conclusions: Registry data is commonly used to report cancer statistics. Registry data may not be accurate for the true incidence of KC since 11.5% were based on radiology alone. Clear cell KC made up 53% of registry diagnosis but 62% on pathology report review. Although papillary and chromophobe incidence did not vary a lot, other types of nonccKC did. This registry data did not differentiate between papillary type I and II. NonccKC should not be considered one entity. One must be aware of the gaps in registry data for KC statistics including overall diagnosis, clear cell and nonccKC subtypes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".