Contradictions, methodological flaws, and potential for misinterpretations in ranking treatments of depression
Bibliographic record
Abstract
In this journal, Malhi et al. recently argue1 that the UK National Institute for Health and Care Excellence (NICE) guidelines for depression2 rank short-term psychodynamic therapy (STPP) as last of 11 therapies recommended for less severe depression and 7 out of 10 treatments recommended for more severe depression. They stress that STPP was ranked by NICE below counseling and that individual cognitive-behavior therapy (CBT) has the highest ranking.1 However, we argue that the NICE guidelines for depression are ambiguous in that they recommend multiple treatments as equal first-line treatments for depression on the one hand, but then go on to rank order them in terms of effectiveness and cost-effectiveness based on the NICE guideline's committee interpretation of the available evidence. We are concerned that this ambiguity leads to possible misinterpretations of the evidence as evidenced by Malhi and colleagues. Furthermore, we briefly summarize several methodological flaws in the NICE guidelines proposed ranking of treatments, which further questions the prioritization of one treatment over another in the treatment of depression. Indeed, in their main document, the NICE guidelines recommend several treatments as first-line treatments for less severe depression, emphasizing that patient preferences and other factors such as experiences with previous treatments are important in deciding the type of treatment that is offered to a given patient2, p. 13, 45: “…. take into account that all treatments in table 1 can be used as first-line treatments” (NICE, 2022, p. 28). Similarly, for more severe depression NICE clearly states that “all treatments in table 2 can be used as first-line treatments” (p. 45). As we will discuss below this is in line with the head-to-head comparisons conducted by NICE themselves as well as independent meta-analytic evidence. Yet, contradicting these recommendations, NICE2, p. 13, 45 also rank orders for these treatments based “on the committee's interpretation of their clinical and cost effectiveness and consideration of implementation factors”. This contradiction creates considerable ambiguity and allows for (mis-)interpretation of the NICE guidelines as prioritizing some treatments over others, as appears to be the case by Malhi et al.1 For example, for less severe depression, Malhi et al. state1, p. 468: “Notably, antidepressants are ranked below CBT, BA, and IPT but are recommended ahead of STPP.” For more severe depression, Malhi et al. concluded from the NICE treatment ranking1, p. 468: “This suggests that individuals with severe acute depression should be offered CBT, BA, antidepressant, individual problem-solving, and counseling prior to considering STPP…”. According to the NICE recommendations, however, treatment ranking is secondary to patient preferences and other factors when it comes to deciding which first-line treatment to provide.2, p. 13, 45 For example, NICE emphasizes2, p. 44–45: “Discuss treatment options with people who have a new episode of more severe depression, and match their choice of treatment to their clinical needs and preferences… use table 2 and the visual summary to guide and inform the conversation [and] take into account that all treatments in table 2 can be used as first-line treatments”. This is consistent with the fact that few statistically or clinically significant differences were found in head-to-head comparisons of the treatments listed by NICE as first-line as discussed below. In their reading of the NICE guidelines, Malhi and colleagues do not mention the emphasis in the NICE guidelines on patient preference and other factors and that NICE stresses that all the listed treatments can be offered as a first-line treatment. In summary, neither the indirect nor direct comparisons carried out by the NICE committee nor the cost-effectiveness analyses or independent research5 support Malhi et al.'s claim of superiority of counseling over STPP, prioritizing CBT over other treatments, and ranking STPP among the least effective treatments.1 In conclusion, we argue that the NICE guidelines for depression are ambiguous and even contradictory. This ambiguity may easily lead to misinterpretations as done by Malhi and colleagues resulting in a misrepresentation of the evidence for psychodynamic psychotherapy and other types of psychotherapy. Furthermore, we highlighted several methodological flaws in the NICE treatment ranking. Presently it is not clear which patients benefit from which empirically-supported treatment. Thus, we continue to discourage the devaluing of efficacious treatments so that as many patients as possible may benefit from them. The following authors have been trained in PDT: FL, AA, PL, and CS. SR has been trained in CBT but has mainly done research on psychodynamic therapy. NH is presently in training of PDT. PL received royalties from Guilford Press, Wiley, Routledge, and Cambridge University Press. AA received royalties from Seven Leaves Press. FL received royalties from Hogrefe Publisher. Data are available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".