A convolutional neural network for high throughput screening of femoral stem taper corrosion
Bibliographic record
Abstract
Corrosion at the modular head-neck taper interface of total and hemiarthroplasty hip implants (trunnionosis) is a cause of implant failure and clinical concern. The Goldberg corrosion scoring method is considered the gold standard for observing trunnionosis, but it is labor-intensive to perform. This limits the quantity of implants retrieval studies typically analyze. Machine learning, particularly convolutional neural networks, have been used in various medical imaging applications and corrosion detection applications to help reduce repetitive and tedious image identification tasks. 725 retrieved modular femoral stem arthroplasty devices had their trunnion imaged in four positions and scored by an observer. A convolutional neural network was designed and trained from scratch using the images. There were four classes, each representing one of the established Goldberg corrosion classes. The composition of the classes were as follows: class 1 ( n = 1228), class 2 ( n = 1225), class 3 ( n = 335), and class 4 ( n = 102). The convolutional neural network utilized a single convolutional layer and RGB coloring. The convolutional neural network was able to distinguish no and mild corrosion (classes 1 and 2) from moderate and severe corrosion (classes 3 and 4) with an accuracy of 98.32%, a class 1 and 2 sensitivity of 0.9881, a class 3 and 4 sensitivity of 0.9556 and an area under the curve of 0.9740. This convolutional neural network may be used as a screening tool to identify retrieved modular hip arthroplasty device trunnions for further study and the presence of moderate and severe corrosion with high reliability, reducing the burden on skilled observers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".