Clinical Trials in Orthopaedics Research. Part II. Prioritization for Randomized Controlled Clinical Trials *
Bibliographic record
Abstract
The American Academy of Orthopaedic Surgeons (AAOS) and Orthopaedic Research Society (ORS) Clinical Trials in Orthopaedics Research Symposium had three major themes including barriers to performing clinical trials, methodology of clinical trials, and prioritization of clinical questions in orthopaedics to address in randomized controlled trials. This paper addresses the latter theme. Clinical experts from the major orthopaedic specialties provided presentations on key issues in their respective fields that had been addressed with clinical trials and the clinical questions that were most appropriate and pressing for clinical trials. Clinicians caring for patients with musculoskeletal disorders must make treatment decisions daily. While the number of clinical trials has increased dramatically in the last two decades, many clinical decisions remain guided by a weak evidence base. The most rigorous evidence for evaluating the efficacy of interventions comes from randomized controlled trials. While it is tempting to envision filling each gap in the clinical evidence base with a randomized controlled trial, trials are resource intensive. The scientific community cannot perform a randomized controlled trial to answer every clinical question. How, then, do clinical scientists prioritize potential trials to select those that yield the most favorable balance between the resources required to conduct the trial and the value of the information a trial would yield? We suggest approaching these critical questions through three complementary lines of inquiry. The first is evaluation of the information that would be gained if the trial were executed successfully. The second is the feasibility of the trial. The final consideration is the resource cost of performing the trial. We ask: What would be gained from the trial? What would it take for the trial to succeed in answering the study question? What would it cost? The first goal of this paper is to articulate a framework for prioritizing clinical trials using these three considerations. The major features of the framework are outlined in Table I. The second goal of the paper is to evaluate the randomized controlled trial topics proposed by speakers at the AAOS-ORS Clinical Trials in Orthopaedics Research Symposium through the lens of this conceptual framework. These topics are summarized in Table II. TABLE I - Schema for Prioritization of Randomized Controlled Trials Value of information gained Prevalence of the condition and its effect on health status Uncertainty regarding efficacy and appropriate indications Variation in practice patterns among clinicians Preliminary evidence of efficacy Enduring nature of the intervention Feasibility Availability of adequate sample size Multiple centers Equipoise (community, individual clinician, and patient) Reliable, valid, responsive measures of outcome Pilot data to ensure protocol is realistic Resource cost of trial Research personnel Space and equipment Management of potential conflicts Assessment of trial costs as weighed against value of information TABLE II - Randomized Controlled Trial Ideas Proposed by Symposium Speakers Treatments to Be Compared According to Specialty Hand and upper extremity Arthroscopic vs. open rotator cuff surgery for rotator cuff tear Proximal interphalangeal joint or wrist arthroplasty vs. nonoperative therapy for advanced arthritis Higher vs. lower dose of injectate for steroid injection of carpal tunnel, shoulder, etc. Efficacy of cognitive behavioral therapy vs. usual care* for most musculoskeletal conditions Sports† Use of single-row vs. double-row anchors in rotator cuff surgery Surgery vs. nonoperative care for rotator cuff tendinopathy Use of autograft vs. allograft in ACL reconstruction Use of single-bundle vs. double-bundle tendon grafts in ACL reconstruction ACL reconstruction vs. nonoperative therapy for symptomatic tear of the ACL Partial resection of the meniscus vs. specific exercise-based regimen in patients with knee osteoarthritis Pediatrics Sports injury prevention program vs. control program to reduce the prevalence of sports injuries Surgical vs. nonoperative management of clavicular fractures in adolescents Trial of growth modulation techniques vs. usual care in spinal surgery Early vs. delayed reduction of dislocated hips in developmental dysplasia of the hip Gait analysis vs. usual process for decision making in cerebral palsy Surgical vs. nonoperative intervention for upper extremity impairment in cerebral palsy Spine‡ Biologic anti-inflammatory medications vs. usual care for sciatica due to disc protrusion BMP growth products vs. usual strategy for spine fusion Lower extremity reconstruction Cemented vs. uncemented fixation in total hip replacement for selected patient groups (e.g., older patients) Ceramic vs. metal heads in total hip replacement Large head diameters (>32 mm) vs. smaller heads in total hip replacement: effect on dislocation and function Cemented vs. uncemented fixation in total knee arthroplasty Patellar resurfacing vs. nonresurfacing in total knee arthroplasty All polyethylene vs. metal-backed tibial components in total knee arthroplasty Cross-linked vs. conventional polyethylene in total knee arthroplasty Unicompartmental vs. total knee arthroplasty in elderly patients with primarily unicompartmental disease Antibiotic-loaded cement vs. usual cement for prophylaxis against infection in total knee arthroplasty Foot and ankle Anticoagulation vs. placebo in foot surgery Diabetic ulcer healing: wound adjuvant vs. boot braces Total ankle replacement vs. arthrodesis for advanced ankle arthritis Toe replacement vs. nonoperative care for advanced toe arthritis Trauma Bone-healing adjuvants vs. usual care for complex fractures Ultrasound stimulation vs. usual care for healing of tibial fractures Wound closure vs. usual care for open extremity fractures Distraction osteogenesis vs. usual care for fracture-healing Total hip arthroplasty vs. hemiarthroplasty in displaced hip fracture Oncology Optimal anticoagulation approaches in patients undergoing orthopaedic oncologic procedures Gene therapy Interleukin-1-based intra-articular gene therapy vs. placebo injection for knee osteoarthritis *As per text, so-called usual care should be specified algorithmically to permit meaningful inference.†ACL = anterior cruciate ligament.‡BMP = bone morphogenetic protein. Value of Information Gained If the trial is executed successfully, will the results be useful to clinicians and patients in their decision to offer a particular treatment? How useful? Does the value of the information gained justify the cost of obtaining it? Several factors influence the value of information to clinicians and patients. Burden Is the problem common and disabling? Trials that address rare conditions, or conditions that have trivial effects on patient well-being, will have limited impact at the societal level. We do not suggest that such trials are unimportant or that they should not be done. In fact, we readily acknowledge that trials that address rare problems are critical for the treatment of those conditions to advance. However, these trials will have less public health impact than trials addressing more prevalent and disabling problems, and typically will have a lower priority. It is noteworthy that randomized controlled trials of rare conditions often pose distinctive logistical challenges to obtain an adequate enrollment. Relevance Over Time A trial comparing two devices that will both be obsolete before the trial is completed may not have an enduring message. In contrast, trials comparing two strategies that remain relevant to practice even as the specific technologies change will continue to guide clinical decisions over time. For example, a trial of lumbar fusion versus laminectomy is more likely to influence practice fundamentally over time than a trial of two competing pedicle screw and plate constructs, neither of which is likely to be used in practice for more than a few years. We recognize that there is an important role for trials that compare competing devices for the same indication. Without such trials, choices between competing devices are made on the basis of nonscientific considerations. However, we suggest that, ideally, these trials should be done before the device enters the marketplace or soon thereafter. We also note that such trials should ideally gather appropriate data to support cost-effectiveness analyses, as cost-effectiveness is an important consideration in assessing competing therapies. Specificity of the Intervention Are the alternative treatment strategies sufficiently specified to be informative? A trial of a well-defined surgical procedure versus a well-defined exercise protocol (e.g., for rotator cuff tear, meniscal tear, or lumbar disc protrusion) will be more interpretable than trials of “surgery” versus “usual nonoperative care,” for which the actual content of surgery and usual nonoperative care varies widely among trial participants. There is a role for comparisons between new treatments and usual care, as most new interventions must compete with the usual way of treating patients. However, investigators ideally should specify and implement usual care in a standardized manner that permits meaningful inference. In fact, detailed delineation of the intervention is a key element of study design. Consider, for example, a trial of surgery versus usual care in which some of the usual care group receive exercise, others receive corticosteroid injections, some receive nonsteroidal anti-inflammatory drugs (NSAIDs), others receive behavioral cognitive therapy, and many receive various combinations of these approaches. If these diverse treatments are offered haphazardly, without a protocol specifying which treatments should be offered and when, the implications of the trial findings for clinical decision making will be unclear. Uncertainty Practice variation provides empirical evidence of uncertainty regarding optimal treatment. If there is no practice variation in the treatment of a particular condition, the clinician community likely will not enroll patients. More importantly, the clinician community might not accept the findings of a trial, even if it shows that the commonly performed treatment is less efficacious than an alternative. On the other hand, if there is considerable variation in treatments used across providers, reflecting uncertainty about appropriate indications and outcomes, a trial is more likely to change practice. The wide variety of approaches to anticoagulation following knee replacement (warfarin, aspirin, heparin, mechanical compression, and combinations of these) offers strong testimony to the uncertainty in the community regarding the optimal anticoagulation approach. We suggest that this level of practice variation and underlying uncertainty makes the clinical community ready to accept a trial. It would be important in such a trial to understand the minimal clinically important difference in event rates and to power the trial accordingly. Preliminary Evidence of Efficacy The case for a randomized controlled trial of an emerging treatment is typically more compelling when there is preliminary evidence of efficacy from smaller controlled studies (strong preliminary evidence), observational studies (less strong), or case series (weak preliminary evidence). If preliminary data do not suggest that the agent (drug, device, or behavioral intervention) is likely to serve as an effective treatment strategy, investment in a randomized controlled trial may not move the field forward. Role of Negative Trials Negative trials are extremely useful. If a more expensive or more harmful intervention is used routinely without rigorous evidence of superiority over less expensive or toxic alternatives, a negative trial can change practice. A striking example in the orthopaedic community is arthroscopic debridement and lavage for osteoarthritis. Following the popularization of arthroscopy in the 1980s, arthroscopic lavage and debridement were performed frequently for osteoarthritis, with some support from small nonrandomized studies. However, two high-quality randomized controlled trials performed in the past decade established that arthroscopic lavage and debridement are no more useful than sham surgery or nonoperative therapy for symptoms of osteoarthritis1,2. This conclusion has been integrated into practice guidelines. “Negative” trials must be interpreted carefully. If a trial is powered to detect the minimal clinically important difference in event rates (or in mean outcome scores) and the finding of the trial is that the difference between groups is less than the minimal clinically important difference, then the trial is negative and the null hypothesis can be rejected. However, if the trial does observe a group difference equal to or greater than the minimal clinically important difference, but the difference fails to reach significance, then the trial must be regarded as inconclusive and further research is needed. Feasibility The critical issue here is whether the trial will be executed successfully once it is initiated. There are four particularly important aspects of trial feasibility. Availability of Eligible Subjects Even if the condition under study is prevalent, the pool of eligible patients may be small. For example, an investigator could readily identify persons with advanced osteoarthritis by enrolling patients in the offices of orthopaedic surgeons who specialize in joint arthroplasty. Persons with early osteoarthritis are more difficult to identify because they are less frequently referred to specialists. Attempts to identify these populations can result in bias if not done carefully. Limitations in the pool of eligible subjects at a single institution may necessitate a multicenter trial, which requires substantially more trial infrastructure, coordination, cost, and more sophisticated statistical considerations than a single-center trial. Outcome Measures Every trial requires a reliable, valid, and responsive primary outcome measure. In many conditions, such as cancer and cardiovascular disease, death and salient clinical events—such as myocardial infarction—typically constitute the primary outcomes. However, orthopaedic procedures are often performed to reduce pain or improve function. These domains are more precisely measured with multi-item scales than with single discrete questions3. While the psychometric properties of such outcome measures are beyond the scope of this article, it is important to recognize that an appropriate validated outcome measure must be selected before a trial can proceed. This field is now rather mature, and an appropriate measure often exists for most orthopaedic conditions and interventions. If an appropriate measure does not exist or has not been validated for a particular condition (as may occur in rarer, less studied conditions), a measure must be developed and/or validated before a trial is implemented. Equipoise The trial will not enroll adequate numbers of patients successfully if potential subjects and enrolling physicians are not comfortable with both treatment options under study in a randomized controlled trial. The community of clinical scientists involved in the trial defines eligibility criteria in order to include only the patients for whom there is genuine uncertainty about the appropriate management. This consensus on eligibility criteria among clinical scientists is called clinical equipoise, or community equipoise4. However, the participating physician investigator who evaluates an eligible patient may have a strong clinical intuition about which treatment strategy will be most helpful. Despite the fact that the patient is eligible, this clinician may be uncomfortable randomizing. This situation illustrates the tension that may exist between community equipoise and individual equipoise. Community equipoise reflects the judgment of a group of well-informed clinicians, typically based on a critical evaluation of the research literature. Individual equipoise, on the other hand, is the state in which the individual clinician is comfortable with both alternatives. If the clinician wishes to exercise clinical judgment rather than recommend randomization to a patient, despite the lack of evidence supporting a particular treatment for that particular patient, then community and individual equipoise collide5,6. This can dampen enrollment of a randomized controlled trial and also introduce bias and limitations in generalizability, since certain patients who are eligible will not be randomized. Of course, patients must also experience equipoise—comfort with both options under study—in order to enroll in a randomized controlled trial. If the vast majority of patients have strong preferences for one treatment or the other, the trial will not succeed. In many surgical trials, 20% to 30% of eligible patients enroll, with the remainder either not referred (because of lack of surgeon equipoise) or referred but not enrolled (because of lack of patient equipoise). Thus, individual surgeon equipoise and individual patient equipoise are required for a successful enrollment. Subjects who ultimately enroll in the trial should be compared carefully with those who are eligible but do not enroll in order to be able to assess generalizability of the enrolled sample to the pool of eligible patients to whom inferences will be made5. Pilot Data Nothing is more reassuring about the feasibility of a research protocol than direct evidence that all aspects of the protocol can be implemented successfully. Pilot data can also provide key parameters necessary for trial design. For example, sample size calculations require an estimate of the effect of the intervention and the of outcome with the intervention and the control with an estimate of the clinically important difference at the individual and group Pilot studies permit of these parameters in the of We note that, even when a trial a with is required to perform the trial. a must be with a research Trials are typically The of randomized controlled trials Trials often require direct patient costs such as and Trials may provide information if such as and studies are in the A multicenter trial requires in the research in each data management a strong and These various costs should be as as in the trial in order to trials are they require trials of and are by the that the trials have potential for bias if the is to influence the and particularly the analysis and of the trial such potential conflicts of and an appropriate between the and the is These conflicts may also trials, particularly if they are by a with a in the trial The potential for bias in studies the of of potential conflicts of While trials are they also provide scientific evidence of A trial of a important can change the of management for the condition Thus, the costs are are the potential Trial Proposed at the AAOS-ORS Clinical Trials Symposium orthopaedic surgeons from the major orthopaedic made presentations at the that key questions in each appropriate and for randomized controlled trials. of these trial are in Table II. The should not be interpreted as the of important trial but rather as a sample that illustrates some potential of these questions have been addressed with trials but would from further with randomized controlled trials. While the topics in Table II were proposed as potential subjects for clinical trials, we note that the randomized controlled trial is not the only for assessing many of these in those in which preliminary data are or observational can be and yield of effect to the sample size for the randomized trial. studies are particularly if the goal of the research is not to treatment efficacy but rather to evaluate the of patients who take a particular from the randomized controlled trial, the observational the for by indication. from the of prioritization criteria in Table many of these trial some but not all of the We provide a few A trial of the dose of injectate for of the rotator cuff or carpal would likely be by the clinical individual enrolling and patients. It would also address an with considerable practice the effect on the health status of the is likely to be A study of unicompartmental versus total knee arthroplasty in select patients would to an important but may be limited by a of surgeons to in this The trial would require a of in order to detect in and The to enroll a and for to a decade or more feasibility A of dislocation following of versus smaller heads in total hip arthroplasty would address an and patients would be to The sample size are because dislocation is an The trial would require centers and Trials of versus nonoperative therapy often address important clinical patients and physicians may have strong about the of even in the of rigorous evidence one strategy over Thus, the and for the typically a major for trials of versus nonoperative such as surgical versus therapy management of rotator cuff a these trials will to many eligible patients to enrollment with strong a preferences are most likely to the In contrast, trials of two surgical techniques are less likely to questions of equipoise. may have no a about single versus double-bundle anterior cruciate reconstruction or versus polyethylene components for total knee arthroplasty. On the other hand, these trials may compare two that will both be by the time of the to or not to likely remain relevant for years. In this the framework we have proposed can be used to the of randomized controlled trial. trial will have one or more features that a and others that do the involved in whether the trial (e.g., and will to and balance these to at prioritization We that the offered in the paper investigators and other to among trial both at the and in making the final decisions about which trials to move for consideration of and
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.876 | 0.976 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.062 | 0.016 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".