Predictors of Response to Cyclo-Oxygenase-2 Inhibitors in Osteoarthritis: Pooled Results from Two Identical Trials Comparing Etoricoxib, Celecoxib, and Placebo
Bibliographic record
Abstract
OBJECTIVE: Nonsteroidal anti-inflammatory drug (NSAID) responses in osteoarthritis (OA) are highly variable, often requiring multiple medication changes. We sought to determine pre-randomization predictors of response to NSAIDs in OA. METHODS: Data were pooled from two identical 26-week double-blind, randomized flare design trials comparing etoricoxib 30 mg/day (N=475), celecoxib 200 mg/day (N=488), and placebo (N=244) in patients with OA of the hip or knee. This analysis was limited to the 12-week placebo-controlled period. Response at Week 12 was defined using Outcome Measures in Rheumatology Clinical Trials and Osteoarthritis Research Society International (OMERACT-OARSI) criteria. Factors were analyzed using logistic regression and included age, race, gender, body mass index, index joint, screening (pre-washout) and baseline (post-washout) Western Ontario and McMaster Universities (WOMAC) Osteoarthritis Index pain, physical function, and stiffness, and patient global assessment of disease status, prior NSAID/coxib or acetaminophen use, American Rheumatology Association functional class, and disease duration. RESULTS: We found that screening WOMAC physical function was the only factor that predicted response in all treatment groups; worse function was associated with lower odds of achieving an OMERACT-OARSI response at 12 weeks (odds ratio 0.84 placebo; 0.87 etoricoxib; 0.89 celecoxib; P<0.05 for all). However, the differences in WOMAC physical function between responders and nonresponders were small (∼5 mm on a 100-mm scale). No factor discriminated between the ability to predict placebo response from active treatment response. CONCLUSIONS: Lower levels of physical function decreased the odds of a response to NSAID treatment in OA, although the clinical significance is unknown given the small differences between responders and nonresponders. No other measured baseline variables consistently predicted response in these studies, which may reflect the known individual variability in NSAID response.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".