Bibliographic record
Abstract
Commentary Over the past half century, total knee arthroplasty (TKA) has evolved into a reliable, cost-effective procedure that provides predictable long-term implant survivorship and meaningful improvements in pain and function in the majority of patients. Additionally, as care pathways have evolved, the health economic burden of individual procedures has decreased1. Given these successes, as well as the marked increases in the prevalence of symptomatic knee osteoarthritis in the United States and elsewhere, it is perhaps not surprising that the incidence of knee arthroplasty is approaching 1 million cases per year in the U.S. alone2. Despite these accomplishments, the outcomes of TKA remain variable. While the majority of patients have a successful outcome, data suggest that as many as 20% of patients have unsatisfactory results with surgery3. At the same time, other patients who are likely to benefit from this procedure may have difficulty accessing it. For example, certain racial groups remain consistently underrepresented among patients undergoing TKA4. With these considerations in mind, Ghomrawi et al. sought to gain a better understanding of the appropriateness of utilization of TKA in the United States with use of data from 2 longitudinal cohorts of patients with risk factors for knee osteoarthritis. Using the Escobar criteria for appropriateness, they found that of 3,123 knees deemed potentially appropriate for TKA, only 9% had undergone TKA within 2 years. Conversely, of all of the patients who had undergone TKA during the study period, 26% did not meet appropriateness criteria and were deemed premature. The authors additionally reported associations between black race and decreased likelihood of having undergone TKA despite meeting appropriateness criteria, as well as associations between elevated body mass index and depression and an increased likelihood of premature TKA. With 91% of potentially appropriate knees in the cohort having not undergone TKA, this study highlights the large societal burden of symptomatic knee osteoarthritis, and suggests that the large number of TKAs performed every year represent only a small proportion of patients living with symptomatic knee osteoarthritis. With this in mind, one wonders how osteoarthritis is affecting the lives of the 91% of “potentially appropriate” patients who did not undergo TKA in this study. Is this silent majority of “potentially appropriate” patients successfully managing their symptoms and maintaining acceptable quality of life through less-invasive treatments such as neuromuscular exercise, activity modification, and non-narcotic oral analgesia? Or do these findings indicate that many individuals are suffering but are either unable or unwilling to undergo arthroplasty surgery? I imagine that both of these possibilities are true and hope that this study will motivate additional work that will lead to a better understanding of these findings. An important limitation of this study rests with the appropriateness criteria used. The Escobar criterwia were developed with use of Delphi methodology by a panel of exclusively musculoskeletal physicians (dominated by orthopaedic surgeons). Perhaps for this reason, the criteria rely largely on physician-assessed factors such as age, radiographic grading, and objective knee stability. However, many adult reconstruction surgeons have learned through experience the limited ability of these objective measures alone to reliably predict a satisfied patient. The literature is increasingly recognizing the importance of factors such as medical comorbidities, patient preferences and values, functional impact, behavioral factors (including patient engagement in the management of their own health), and patient experience with other treatment options in informing the appropriateness of TKA surgery as well as in predicting surgical outcomes from a patient perspective. With this in mind, the criteria used may not accurately reflect contemporary decision-making around TKA. We also should consider the fact that appropriateness for a particular intervention is not a binary measure. For knee osteoarthritis in particular, there are a range of invasive and noninvasive treatment options available, many of which could be potentially appropriate for a given patient. Helping patients to identify the best option for them at a given point in time requires a condition-centered approach that considers the full range of other treatment options available, including the expected risks and benefits of each for that particular patient. Thus, even though a particular treatment may be “potentially appropriate,” it might not be the most appropriate at that time. Notwithstanding the limitations of the study, the authors should be commended for exploring the question of appropriateness of contemporary total knee arthroplasty, specifically, the extent to which this procedure might be both underused and overused. While the answers supplied by the study are limited, they provide an opportunity to reflect on how our understanding of the appropriateness of TKA has evolved over time from a largely surgeon-centered perspective. They also highlight the work that must still be done to better inform ongoing transitions toward integrated, high-value, condition-centered care, and to ensure that every patient with symptomatic knee osteoarthritis receives advice and treatment that is most appropriate for their individual circumstances, preferences, and goals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.147 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.005 | 0.009 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.007 | 0.003 |
| Research integrity | 0.034 | 0.038 |
| Insufficient payload (model declined to judge) | 0.029 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".