Global Burden of Disease estimates of depression – how reliable is the epidemiological evidence?
Bibliographic record
Abstract
OBJECTIVES: To re-assess the quality of the epidemiological studies used to estimate the global burden of depression 2000, as published in the GBDep study. DESIGN: Primary and secondary data sources used in the global burden of depression estimate were identified and assigned to country of origin. Each source was assessed with respect to completeness and representativeness for national/regional estimates and against the inclusion criteria used by the scientific team estimating GBDep. SETTING: Not applicable. PARTICIPANTS: Not applicable. MAIN OUTCOME MEASURES: Not applicable. RESULTS: First, National estimates: The 28 scientific sources cited in the GBDep study related to 40 of the 191 WHO member countries. The EURO region had studies relating to 15 of 52 countries whereas AFRO region had studies for only three of 46 countries. Only six of the 40 countries had data drawn from a nationally representative population: the three AFRO country studies were based on a single village or town and, likewise, SEARO region had no nationally representative data; second, GBDep criteria: GBDep inclusion criteria required study sample size of more than 1000 people; 19 (45%) of the 42 studies did not meet this criterion. Sixteen (44%) of 36 studies did not meet the requirement that studies show a clear sample frame and method. GBD estimates rely on estimates of incidence; only two of the 42 country studies provided incidence data (Canada and Norway), the remaining 34 studies were prevalence studies. Duration of depression is based on three studies conducted in the USA and Holland. CONCLUSIONS: Most studies exhibit significant shortcomings and limitations with respect to study design and analysis and compliance with GBDep inclusion criteria. Poor quality data limit the interpretation and validity of global burden of depression estimates. The uncritical application of these estimates to international healthcare policy-making could divert scarce resources from other public healthcare priorities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".