19. Prevalence of diagnostic discordance: A retrospective analysis of autopsy findings and clinical diagnoses
Bibliographic record
Abstract
The prevalence of medical error in health care systems has compromised the quality of health care delivery. The research on medical errors in hospitalized population has consistently revealed high rates of misdiagnosis. Autopsy examination has been an established tool for quality assurance programs. The objective of this study was to determine the discrepancy rates between clinical and autopsy findings in patients admitted to various hospitals (Royal University hospital, RUH; St. Paul’s hospital, SPH; Saskatoon city hospital, SCH) of Saskatoon Health Region. A retrospective record review of the medical and autopsy charts was carried out for all the deceased adult in-patients admitted during the years 2002, 2003, and 2004. All autopsies were carried out either in the morgue of RUH or at SPH hospital. A total of 3416 in-patient deaths were registered during the study period. Autopsies were performed on 206 of the deceased resulting in an autopsy rate of 6%. In accordance with selection criteria, 158 cases were included for this study. The mean age of subjects was 66.6 ± 15.2 years with a range of 16 – 94 years. The total concordance rate in this study between clinical and autopsy diagnoses was 75.3%. The discordance rate was 20.9% and in 3.8% of the study population a conclusive clinical or autopsy diagnoses was not finalized. The concordance and discordance rate between clinical diagnoses and autopsy findings when compared between the patients of two hospitals (RUH and SPH) were not significantly different. These results suggest that despite the technical advances in medical and diagnostic modalities, there still persists a significant discordance in clinical and autopsy diagnoses. Our study confirms the wide prevalence of diagnostic discrepancies in the health care system and emphasizes the value of autopsy as an effective quality improvement and educational tool with a strong impact on quality management. 
 Roosen J, Frans E, Wilmer A, Knockaert DC, Bobbaers H. Comparison of premortem clinical diagnoses in critically ill patients and subsequent autopsy findings. Mayo Clin Proc 2000; 75:562-67. 
 Sonderegger-Iseli K, Burger S, Muntwyler J, Salomon F. Diagnostic errors in three medical eras: a necropsy study. Lancet 2000; 10355:2027-31.
 Kalra J. Medical Error: An Introduction to Concepts. Clinical Biochemistry 2004;37:1043-51.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.040 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.016 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".