Hip Resurfacing versus Total Hip Arthroplasty: A Systematic Review Comparing Standardized Outcomes
Bibliographic record
Abstract
BACKGROUND: Metal-on-metal hip resurfacing was developed for younger, active patients as an alternative to THA, but it remains controversial. Study heterogeneity, inconsistent outcome definitions, and unstandardized outcome measures challenge our ability to compare arthroplasty outcomes studies. QUESTIONS/PURPOSES: We asked how early revisions or reoperations (within 5 years of surgery) and overall revisions, adverse events, and postoperative component malalignment compare among studies of metal-on-metal hip resurfacing with THA among patients with hip osteoarthritis. Secondarily, we compared the revision frequency identified in the systematic review with revisions reported in four major joint replacement registries. METHODS: We conducted a systematic review of English language studies published after 1996. Adverse events of interest included rates of early failure, time to revision, revision, reoperation, dislocation, infection/sepsis, femoral neck fracture, mortality, and postoperative component alignment. Revision rates were compared with those from four national joint replacement registries. Results were reported as adverse event rates per 1000 person-years stratified by device market status (in use and discontinued). Comparisons between event rates of metal-on-metal hip resurfacing and THA are made using a quasilikelihood generalized linear model. We identified 7421 abstracts, screened and reviewed 384 full-text articles, and included 236. The most common study designs were prospective cohort studies (46.6%; n = 110) and retrospective studies (36%; n = 85). Few randomized controlled trials were included (7.2%; n = 17). RESULTS: The average time to revision was 3.0 years for metal-on-metal hip resurfacing (95% CI, 2.95-3.1) versus 7.8 for THA (95% CI, 7.2-8.3). For all devices, revisions and reoperations were more frequent with metal-on-metal hip resurfacing than THA based on point estimates and CIs: 10.7 (95% CI, 10.1-11.3) versus 7.1 (95% CI, 6.7-7.6; p = 0.068), and 7.9 (95% CI, 5.4-11.3) versus 1.8 (95% CI, 1.3-2.2; p = 0.084) per 1000 person-years, respectively. This difference was consistent with three of four national joint replacement registries, but overall national joint replacement registries revision rates were lower than those reported in the literature. Dislocations were more frequent with THA than metal-on-metal hip resurfacing: 4.4 (95% CI, 4.2-4.6) versus 0.9 (95% CI, 0.6-1.2; p = 0.008) per 1000 person-years, respectively. Adverse event rates change when discontinued devices were included. CONCLUSIONS: Revisions and reoperations are more frequent and occur earlier with metal-on-metal hip resurfacing, except when discontinued devices are removed from the analyses. Results from the literature may be misleading without consistent definitions, standardized outcome metrics, and accounting for device market status. This is important when clinicians are assessing and communicating patient risk and when selecting which device is most appropriate for individual patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.012 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.002 | 0.007 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".