Factors Associated with Physician Agreement on Verbal Autopsy of over 11500 Injury Deaths in India
Bibliographic record
Abstract
INTRODUCTION: Worldwide, injuries account for 9.8% of all deaths. The majority of these deaths occur in low- and middle-income countries where vital registration systems are often inadequate. Verbal autopsy (VA) is a tool used to ascertain cause of death in such settings. Validation studies for VA using hospital diagnosed causes of death as comparisons have shown that injury deaths can be reliably diagnosed by VA. However, no study has assessed the factors that may affect physicians' abilities to code specific causes of injury death using VA. METHOD/PRINCIPAL FINDINGS: This study used data from over 11,500 verbal autopsies of injury deaths from the Million Death Study (MDS) in which 6.3 million people in India were monitored from 2001-2003 for vital events. Deaths that occurred in the MDS were coded by two independent physicians. This study focused on whether physician agreement on the classification of injury deaths was affected by characteristics of the deceased and respondent. Agreement was analyzed using three primary methods: 1) kappa statistic; 2) sensitivity and specificity analysis using the final VA diagnosed category of injury death as gold standard; and 3) multivariate logistic regression using a conceptual hierarchical model. The overall agreement for all injury deaths was 77.9% with a kappa of 0.74 (99% CI 0.74-0.75). Deaths in the injury categories of "transport", "falls", "drowning" and "other unintentional injury" occurring outside the home were associated with greater physician agreement than those occurring at home. In contrast, self-inflicted injury deaths that occurred outside the home were associated with lower physician agreement. CONCLUSIONS/SIGNIFICANCE: With few exceptions, most characteristics of the deceased and the respondent did not influence physician agreement on the classification of injury deaths. Physician training and continued adaptation of the VA tool should focus on the reasons these factors influenced physician agreement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".