Predicting placebo analgesia in patients with chronic pain using natural language processing: a preliminary validation study
Bibliographic record
Abstract
ABSTRACT: Patients with chronic pain show large placebo effects in clinical trials, and inert pills can lead to clinically meaningful analgesia that can last from days to weeks. Whether the placebo response can be predicted reliably, and how to best predict it, is still unknown. We have shown previously that placebo responders can be identified through the language content of patients because they speak about their life, and their pain, after a placebo treatment. In this study, we examine whether these language properties are present before placebo treatment and are thus predictive of placebo response and whether a placebo prediction model can also dissociate between placebo and drug responders. We report the fine-tuning of a language model built based on a longitudinal treatment study where patients with chronic back pain received a placebo (study 1) and its validation on an independent study where patients received a placebo or drug (study 2). A model built on language features from an exit interview from study 1 was able to predict, a priori, the placebo response of patients in study 2 (area under the curve = 0.71). Furthermore, the model predicted as placebo responders exhibited an average of 30% pain relief from an inert pill, compared with 3% for those predicted as nonresponders. The model was not able to predict who responded to naproxen nor spontaneous recovery in a no-treatment arm, suggesting specificity of the prediction to placebo. Taken together, our initial findings suggest that placebo response is predictable using ecological and quick measures such as language use.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".