The Reliability of Digitized Radiographs for Dental Identification: A Web-Based Study
Bibliographic record
Abstract
In the era of Daubert and other judicial rulings pertaining to the acceptability of forensic evidence, it is increasingly important that experts are able to testify that their methods have been scientifically tested and that error rates and other factors relating to reliability have been published. The purpose of this study was to determine the reliability of digitized radiographic comparisons for the purposes of dental identification. Participants with various forensic backgrounds and experience levels were passively recruited to the website. Ten forensic identification cases composed of antemortem and postmortem dental radiographs were supplied to examiners using a bespoke website. Participants responded to the cases on two occasions after a one-month washout interval using the ABFO conclusion levels for forensic identifications. A total of 115 first attempts and 87 matched second attempts were received. Of the total responses, 72% were dentally trained respondents who had completed at least one forensic identification case; of these, 38% were experienced forensic dentists who had completed more than 25 identifications. Data relating to accuracy, intra- and inter-examiner agreement, and the effect of case difficulty are presented. Mean accuracy was 85.5% for all cases, with the experienced forensic dentists obtaining a 91% success rate. The inter-examiner agreement on the negative identification cases was classified as poor. The data suggest that dental identifications resulting from the comparison of postmortem and antemortem radiographs are valid, accurate, and reliable when undertaken by experienced odontologists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.025 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".