Validating the diagnostic accuracy of medical certification: a population-level comparison between verbal autopsy and Saudi medical records causes of death of deceased with type 2 diabetes
Bibliographic record
Abstract
BACKGROUND: In contexts where certifying causes of death (COD) is inadequate - either in industrialized or non-industrialized countries - verbal autopsy (VA) serves as a practical method for determining probable COD, helping to address gaps in vital data. OBJECTIVE: This study aimed to validate the diagnostic accuracy of medical certifications at a population level by comparing COD obtained from medical records against those derived from VA in Saudi Arabia. METHOD: Death records from 2018 to 2021 were collected from a type 2 diabetes mellitus register at a major specialist hospital in Makkah. Three hundred and two VA interviews were completed with deceased patients' relatives, and the probable COD was determined using InterVA-5 software. Lin's concordance correlation coefficient was applied to examine similarities of the cause-specific mortality fractions (CSMFs) based on International Classification of Diseases chapters from both verbal autopsy causes of death (VACOD) and the physician review causes of death (PRCOD). RESULTS: Overall, the findings demonstrated a moderate level of concordance of COD at the population between VACOD and PRCOD. However, the CSMFs for various COD categories derived from both sources showed a broad spectrum of absolute differences, with the largest disparities observed among the most prevalent COD categories. CONCLUSION: PRCOD was found to overestimate population-level endocrine/metabolic and respiratory disease COD while underestimating circulatory disease, demonstrating medical certification challenges. Conversely, affirming previous literature on prevalent COD in Saudi Arabia, VA appears to deliver a plausible assessment, further strengthening its potential to integrate within the Saudi health system towards an augmented medical certification process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".