CORR Insights®: Projections of Primary TKA and THA in Germany From 2016 Through 2040
Bibliographic record
Abstract
Where Are We Now? Total joint arthroplasty (TJA) of the hip and knee provide truly impressive improvements in health-related quality of life for patients with advanced symptomatic knee and hip osteoarthritis [5]. TJA and spine surgery are also archetypes for “preference-sensitive” care, meaning that the physician and patient have wide discretion over whether to pursue surgery or nonoperative approaches [7]. As such, we know that the use of TJA (measured as procedures per population per year) varies widely from place to place, even among developed countries with similar sociodemographics and healthcare systems [3, 8]. Therefore, unlike procedures such as surgery for hip fractures, where the use or demand is largely fixed and inelastic [4], the use of TJA is thought to be influenced by numerous external factors, including patient preferences and attitudes [2], reimbursement and coverage of TJA by payers, and macrolevel factors such as national income and wealth [1]. But there is much that we do not know. Rupp et al. [9] projected TJA use rates in Germany between 2016 and 2040. The investigators should be commended on their efforts to project what the demand for TJA might be in Germany during the next 25 years. Some of their assumptions, such as aging of the current population, gender, and birth and death rates, can be projected with reasonable confidence. Other assumptions, including immigration rates or changes in reimbursement policies for TJA, are more difficult to predict. Even with these important limitations, the authors provide important and reasonable estimates of TJA use rates between 2016 and 2040 under the most likely scenarios. Their findings of a 45% increase in TKA volume (55% increase in use) and 23% increase in THA volume (29% increase) are important. Hospitals can use these numbers for planning operating room capacity, and medical training programs can use these estimates to guide training program slots and needs. Equally important, payers can use these numbers to examine whether projected spending levels are sustainable. Where Do We Need To Go? In an ideal world, each country and/or jurisdiction would know two things: the current use of all common and costly procedures paid for by government and accurate projections of the future use of each procedure and its financial impact. With current use numbers in hand, governments would be able to compare the current use of assorted procedures in their own country (for example, “what is our per capita rate of TKA and how does that compare with our rate of spine surgery?”). Equally important, such information would allow for a across-country comparisons of use rates of the same procedure (for example, “what is the United States per capita rate of TKA and how does it compare with that of Germany?”). Such use data would begin to give patients, policy makers, and physicians much-needed answers to fundamental questions about which procedures are being performed too frequently and which are underused or deserve additional investment. With future projections for common and costly procedures, jurisdictions would have the information they need to adapt their workforce and hospitals to adjust their mix of procedures and services in an evidence-based manner. As long as governments are not just “a” but “the” major payer of health care, it seems perplexing that neither current use estimates nor predictions of future growth are routinely made in the United States and elsewhere. How Do We Get There? First, it seems reasonable for all public payers to either develop internal capabilities or fund external consultants to develop a “use scorecard.” This scorecard might look something like the Hospital Compare website maintained by the Center for Medicare and Medicaid Service (https://www.medicare.gov/hospitalcompare/search.html) or The Dartmouth Atlas (https://www.dartmouthatlas.org/). Procedures could be identified for the scorecard based on volume, growth in volume over time, cost to the healthcare system, or perhaps expert opinion. Such data could be used internally for planning purposes by policy makers and disseminated externally to groups such as the World Health Organization for international benchmarking. Second, public payers could develop standardized methods for projecting growth of these procedures during the next 10 to 30 years [6]. Without standardized methods, it is difficult to ascertain whether differences in projections arise from differences in methods or from differences in model inputs (such as spending and population growth). With these data, government payers will finally have robust information about their current status (such as, “what is our use of TKA today?”) and projections for the future (such as, “how many TKAs are we likely to be performing in 2035 if we continue on our current trajectory?”). Without these data, we are truly flying blind.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".