Evaluation of accuracy of deep learning and conventional neural network algorithms in detection of dental implant type using intraoral radiographic images: A systematic review and meta-analysis
Bibliographic record
Abstract
STATEMENT OF PROBLEM: With the growing importance of implant brand detection in clinical practice, the accuracy of machine learning algorithms in implant brand detection has become a subject of research interest. Recent studies have shown promising results for the use of machine learning in implant brand detection. However, despite these promising findings, a comprehensive evaluation of the accuracy of machine learning in implant brand detection is needed. PURPOSE: The purpose of this systematic review and meta-analysis was to assess the accuracy, sensitivity, and specificity of deep learning algorithms in implant brand detection using 2-dimensional images such as from periapical or panoramic radiographs. MATERIAL AND METHODS: Electronic searches were conducted in PubMed, Embase, Scopus, Scopus Secondary, and Web of Science databases. Studies that met the inclusion criteria were assessed for quality using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) tool. Meta-analyses were performed using the random-effects model to estimate the pooled performance measures and the 95% confidence intervals (CIs) using STATA v.17. RESULTS: Thirteen studies were selected for the systematic review, and 3 were used in the meta-analysis. The meta-analysis of the studies found that the overall accuracy of CNN algorithms in detecting dental implants in radiographic images was 95.63%, with a sensitivity of 94.55% and a specificity of 97.91%. The highest reported accuracy was 99.08% for CNN Multitask ResNet152 algorithm, and sensitivity and specificity were 100.00% and 98.70% respectively for the deep CNN (Neuro-T version 2.0.1) algorithm with the Straumann SLActive BLT implant brand. All studies had a low risk of bias. CONCLUSIONS: The highest accuracy and sensitivity were reported in studies using CNN Multitask ResNet152 and deep CNN (Neuro-T version 2.0.1) algorithms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.003 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".