Automated quality evaluation of dental panoramic radiographs using deep learning
Bibliographic record
Abstract
Purpose: Panoramic radiographs are instrumental in dental diagnosis but face quality issues related to contrast, artifacts, positioning, and coverage, which can impact diagnostic accuracy. Although expert assessment is the accepted standard, it is time-consuming and prone to inconsistency. Artificial intelligence offers an automated, objective solution for evaluating radiograph quality, increasing efficiency and reducing inter-rater variability. Materials and Methods: This study aimed to develop a deep learning (DL)-based model for evaluating the quality of dental panoramic radiographs. A dataset of 1,000 panoramic images, collected from 2018 to 2023, was assessed by 2 trained dentists using predefined grading criteria for contrast/density, artifact presence, coverage area, patient positioning, and overall quality. These expert-annotated scores were used as the ground truth to train and validate 5 YOLOv8 classification models, each targeting a specific quality criterion. The models' performance was evaluated on a separate test set using performance metrics. Results: The YOLOv8 models achieved classification accuracies of 87.2%, 74.1%, 77.3%, 97.9%, and 79.3% for artifact detection, coverage area, patient positioning, contrast/density, and overall image quality, respectively. The model used to classify images as clinically acceptable or unacceptable exhibited an average accuracy of 81.4%, demonstrating its potential for real-world application. Conclusion: These findings highlight the feasibility of DL-based automated image quality assessment for panoramic radiographs. The high accuracy of the proposed model suggests its potential integration into clinical workflows to assist practitioners in efficiently evaluating radiograph quality. Additionally, such a model could represent an educational tool for dental students, improving radiographic techniques and reducing unnecessary retakes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".