Letter to the Editor: Is the Survivorship of Birmingham Hip Resurfacing Better Than Selected Conventional Hip Arthroplasties in Men Younger Than 65 Years of Age? A Study from the Australian Orthopaedic Association National Joint Replacement Registry
Bibliographic record
Abstract
To the Editor, This study [4] compares Birmingham hip resurfacing—which is the only hip resurfacing procedure with long-term follow-up—with the best-performing THA implant in the Australian Joint Replacement Registry. Any device compared with the best-performing device will have poorer long-term survival, so the conclusions in this study are not justified by the data. If the authors feel that one must compare the best with the best, because hip resurfacing has a learning curve and poor implant survivorship is associated with steeper cups, the authors should have only compared surgeons with no Birmingham hip resurfacing failure in their first 75 procedures to the best-performing THAs. More importantly, is the Australian registry recommending that all THAs in this age group be performed using only the three selected implant designs? The registry data are flawed by a myriad of confounding variables both known and unknown, including surgeon and patient variables. It is likely that these confounding variables have a greater impact on the implant’s survival than the selection of the device. These confounding variables are not evenly distributed across different devices. By selecting the best-performing device, Stoney et al. [4] may just be selecting the device that coincidentally has the most favorable set of confounding variables. It would have made much more sense to compare the Birmingham hip resurfacing procedure to all THAs in this age group in the Australian Joint Replacement Registry. Could the authors please compare the implant survival data of all THAs with those of the Birmingham hip resurfacing procedure? We will never be able to measure all of these variables, but at least this approach would mitigate the effect of confounding variables. The Australian Joint Replacement Registry showed higher mortality in patients undergoing THA than in patients undergoing hip resurfacing [1]. This difference in mortality is most likely to be a result of confounding variables due to patient selection. However, this is not clear from the data available, and it is theoretically possible that the increased mortality is an effect of the procedure—in other words, amputating the proximal femur and instrumenting the femoral canal somehow causes harm that leads to an increased risk of death compared with preserving the proximal femur and not invading the femoral canal. Studies of other registry data have found that the difference in mortality persists after correcting for multiple confounding variables [2, 3]. Could Stoney et al. [4] please report the difference in mortality between the two groups in their study? These data are available to the registry and could easily have been presented in this report. The authors have identified selection bias as a major limitation of this study that one might argue invalidates their conclusions. Finally, understanding the difference in mortality will help the reader put into perspective the true benefits to the patient and their loved ones of one implant design versus another.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.038 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.011 | 0.014 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".