Attempted emulation of a randomised depression screening trial
Bibliographic record
Abstract
Chen et al.1Chen Y.-L. Wu M.-S. Wang S.-H. et al.Effectiveness of health checkup with depression screening on depression treatment and outcomes in middle-aged and older adults: a target trial emulation study.Lancet Reg Health West Pac. 2024; 43100978Google Scholar used observational data from Taiwan’s Adult Preventive Health Checkup Program to attempt to emulate a depression screening trial by comparing adults who attended a free health checkup to matched counterparts who did not. They reported that people in the health checkup arm were more likely to receive new depression treatment, had lower risk of psychiatric hospitalisation, and that those 65 years and older had higher suicide risk. A depression screening trial should enrol and randomly allocate patients not known to have depression; screen participants in the screening arm but not those in the comparator arm; ensure that participants in both arms have comparable health care access and, if determined to have depression, similar depression care options; and assess achievable screening outcomes, such as depression symptoms or diagnoses.2Thombs B.D. Ziegelstein R.C. Does depression screening improve depression outcomes in primary care?.BMJ. 2014; 348g1253Crossref PubMed Scopus (40) Google Scholar Emulated trials must approximate design elements as closely as possible, including treatment strategies, assignment procedures, and outcomes.3Hernán M.A. Robins J.M. Using big data to emulate a target trial when a randomized trial Is not available.Am J Epidemiol. 2016; 183: 758-764Crossref PubMed Google Scholar Emulating random assignment requires being able to make a strong case that patients in different trial arms are similar except for their treatment assignment.3Hernán M.A. Robins J.M. Using big data to emulate a target trial when a randomized trial Is not available.Am J Epidemiol. 2016; 183: 758-764Crossref PubMed Google Scholar It is unlikely that Chen et al.’s emulated trial arms achieved this. Using a small number of variables to create propensity scores would not likely address confounding from comparing people who sought preventive health care, outside of normal care, to people who did not. Health care available in the two trial arms, beyond depression screening, was not comparable. First, people in the screening arm had (1) health checkups, including a full personal and family history, physical examination, blood tests, and urine tests (2) plus depression screening. People in the comparator arm had neither. Second, people in the screening arm could access additional health checkups and depression screening in the years following the index health checkup, but patients in the comparator arm could not; they were censored if they did. Outcomes did not reflect benefits or harms that would be expected from depression screening or that would normally be included in a depression screening trial. Receiving new treatment occurs with more health care exposure but is not a health benefit. Since depression screening is done to find otherwise undetected cases, trials target symptoms or incident diagnoses. No depression screening trials have targeted hospitalisation and suicide outcomes as in Chen et al.’s study.4Thombs B.D. Markham S. Rice D.B. Ziegelstein R.C. Does depression screening in primary care improve mental health outcomes?.BMJ. 2021; 374: n1661Crossref PubMed Scopus (5) Google Scholar Several well-conducted depression screening trials have reported that screening did not improve mental health outcomes.4Thombs B.D. Markham S. Rice D.B. Ziegelstein R.C. Does depression screening in primary care improve mental health outcomes?.BMJ. 2021; 374: n1661Crossref PubMed Scopus (5) Google Scholar Results reported by Chen et al. do not inform the evidence base further. The authors declare no competing interests. Funding: Dr. Thombs is supported by a Tier 1 Canada Research Chair outside of the present work. There was no funding for the correspondence. Effectiveness of health checkup with depression screening on depression treatment and outcomes in middle-aged and older adults: a target trial emulation studyHealth checkups with depression screening could potentially promote depression treatment and reduce the risk of psychiatric hospitalisation; however, there was no effect on suicide. The treatment rate for depression remained low after screening for depression. Further attention to enhance referral and treatment is required. Full-Text PDF Open Access
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".