Response to commentaries
Bibliographic record
Abstract
Measurement-based care using DSM-5 should be embedded in services now with scope for further development following efficacy and implementation trials. We are grateful to Kathy Bradley and her colleagues 1 and Dennis McCarty 2 for their helpful comments on our article on measurement-based care (MBC) 3. They draw on a wealth of applied experience in efficacy and implementation research in primary care (PC) and substance use disorder (SUD) treatment programmes. Their reflections emphasize the challenge of embedding sustainable clinical measurement of patient progress. Bradley and colleagues illustrate the way PC clinicians’ use of open-ended questions about SUD symptoms can facilitate discussions with patients, beginning with topics and experiences which are important to the person. This is an ideal way to help the patient feel understood, to build therapeutic alliance and review interventions offered. We had this in mind when envisaging DSM-5 criteria as the springboard to a collaborative discussion on the patient's specific SUD-related thoughts, emotions and behaviours. The aim is that treatment goals reflect personal priorities. We agree with Bradley's group that we must ensure the MBC administrative burden is minimized. This has been a priority for all contemporary developers of screening 4 and clinical instruments 5, reflecting a move away from unwieldy scale batteries and a focus on selecting items which provide the most useful information. As a minimum, we think drug craving and substance use should be monitored during the first 2 weeks after medication stabilization and regularly thereafter, to guide decisions as to whether treatment should be maintained, adjusted or switched. Brief items and scales are likely to correlate with DSM-5 SUD criteria but, importantly, they do not confirm remission. McCarty offers insight on his work with providers striving for process improvements in patient retention and engagement 6. He reports how managers stopped asking therapists to complete the brief Session Rating Scale and Outcome Rating Scale 7, 8 after inconsistent use and difficulties with data capture and feedback. This demonstrates how even a very brief measure can flounder in busy clinics. We agree that inconsistent completion confounds validity. Less frequent, consistent completion would be sufficient—and there are also other options. As we discussed, the PHQ-9 is commonly used to monitor symptoms for depression and inform treatment decisions. It is usually self-completed, so it does not compete with session time. Digitally enabled adjunctive psychological therapies can now be offered on-line or through mobile applications to help patients engage in self-monitoring of craving and self-study with discussion face-to-face with a therapist 9, 10. We appreciate that PC providers may query the cost of MBC, but in our view it is clearly warranted by the urgent need for improved treatment outcomes, especially in the context of the OUD epidemic. Clinical trials are needed to study MBC efficacy and implementation. Indicators that guide the way will emerge as a standard—but the objective must always be to help patients access and remain in effective, personalized treatment for as long as is needed. In our view, DSM-5 (or ICD-11) diagnosis and ongoing remission assessment, with interventions matched to symptoms, should be included as the standard of care in all out-patient and PC settings. J.M. declares research funding from Indivior to King's College London for a randomized controlled trial of injectable buprenorphine for the treatment of opioid use disorder.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.026 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".