Seeking a simple measure of analgesia for mega-trials: is a single global assessment good enough?
Bibliographic record
Abstract
We sought to investigate the potential of using a simple global estimation ('How effective do you think the treatment was?') as a measure of efficacy by comparing it with at least 50%maxTOTPAR (at least 50% of the maximum possible pain relief) in acute pain studies. One hundred and fifty randomized, double-blind trials included in 11 systematic reviews of single dose, oral analgesics for postoperative pain were used as a source of data. The relationship between the proportion of patients reporting the top two or three values on a five-point global scale and the proportion with at least 50%maxTOTPAR was investigated. Twenty-six trials provided data on the proportion reporting the top two categories (very good or excellent) and 27 gave data on the top three categories (good, very good or excellent). The relationship between the percentage of patients recording the top two categories on a five-point global scale and the proportion with at least 50%maxTOTPAR was fair (r(2)=0.67). That for the top three categories was less good (r(2)=0.57). Similar numbers-needed-to-treat were calculated for aspirin 600/650 mg and ibuprofen 400 mg using at least 50%maxTOTPAR and the top two categories. No real difference was seen in the correlation for standard wording compared to non-standard wording. Individual patient data were also used from four randomized, placebo-controlled, double-blind trials in postoperative pain. The frequency distribution for %maxTOTPAR was plotted for patients reporting each of the five categories on the global scale. A global assessment provides similar measures of analgesic efficacy as TOTPAR derived from hourly measurements, but the effects of adverse effects have yet to be understood.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.004 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".