Commentary: Explaining enormous variations in rates of disorder in trauma-focused psychiatric epidemiology after major emergencies
Bibliographic record
Abstract
There has been a surge of interest in the last 20 years in the mental health effects of conflict and other major disasters in low- and middle-income countries (LAMIC). In particular, post-traumatic stress disorder (PTSD) and major depression have received substantial attention. It has become evident that there are large, unexplained variations in prevalence rates identified through trauma-focused psychiatric epidemiology in such settings. For example, Mollica et al.'s1 classic study found prevalence rates of PTSD of 15% among genocide-exposed Cambodians, while Neugebauer et al.'s2 sophisticated report in this issue of the Journal identifies rates of 53–62% of PTSD in genocide-exposed Rwandans. The wide variation in the prevalence rates in studies of PTSD and depression may be attributable to differences in context, methodology or both. Discussion sections of reports often highlight only a few factors that could explain the size of obtained rates. Although peer review helps shape discussion sections, authors usually have enormous discretion in deciding what factors to report. Readers are left with the challenge of tracking all reported and unreported methodological and contextual factors that could explain a study's results. We have developed a scheme that may help to systematically identify factors influencing the size of observed prevalence rates of disorders in populations affected by major emergencies in LAMIC. The scheme may prove useful for readers and journal peer reviewers alike. The scheme, which we will apply below to Neugebauer et al.'s2 study, was built as follows. We searched the following medical, psychiatry and speciality journal websites: American Journal of Psychiatry; Archives of General Psychiatry; British Journal of Psychiatry; British Medical Journal; Culture Medicine and Psychiatry; JAMA, Journal of Traumatic Stress, Lancet, Psychological Medicine; Social Science and Medicine; and Transcultural Psychiatry for studies published after 1998 with data collected on depression or PTSD among civilians after major emergencies in LAMIC (references available upon request). Of 43 studies, 11 (26%) pertained to major natural disasters and 32 (74%) pertained to major human-made disasters (e.g. war). All articles were original contributions, and 40 (93%) made comments explaining the magnitude of findings in the articles’ discussion sections. In addition, we reviewed editorials, commentaries and letters to the editor linked to the identified articles. We thematically analysed discussion sections of all papers. We categorized authors’ explanations for observed rates as (i) either methodological or contextual in nature and (ii) explaining either relatively higher or lower observed rates. In addition, we categorized some explanations as (iii) reflecting general methodological limitations causing uncertainty in the validity of the study, with unknown impact on the magnitude of observed rates. Table 1 provides an overview of the explanations for relatively high or low rates ascribed in these studies. Schematic overview of possible influences on uncertainty in observed prevalence rates (with ratings applied to the Neugebauer et al.2 study) Sampling problems Over-sampling of severely affected areas (1) Over-sampling of likely high risk demographic groups (e.g. women) (1) Socially desirable over-reporting (acquiescence) Expectations raised by the way the overall survey and its questions are framed (3) Interviewer's own expectations (3) Help-seeking behaviour among respondents (3) Measures cover life-time prevalence (1) Measures not fully re-validated for specific emergency-affected context To-be-expected distress (e.g. ongoing experiences of loss and misery) not necessarily consistently distinguished from symptoms of mental disorder (4) Reality-based fears (e.g. due to risk of exposure to ongoing further violence) not necessarily consistently distinguished from symptoms of mental disorder (4) Sampling problems Small sample size Non-probability sampling (3) Low response rates (1) Unreported or low reliability: Internal consistency (1) Test–retest reliability (3) Inter-rater reliability (3) Use of incorrect diagnostic criteria (5) Uncertainty whether translation is adequate (1) Unknown or weak criterion validity of exposure measures (1) Unknown or weak criterion validity of outcome measures (3) Uncertainty whether measures concepts are culturally valid (2) Sampling problems Over sampling of less affected areas (1) Under-sampling of likely high-risk groups (1) Underreporting due to problems in recall (1) Socially under-desirable reporting (social stigma) Declarations of anonymity/confidentiality not persuasive. (3) Lack of privacy in interview settings (3) Measures cover point prevalence (1) Trauma-induced mental disorder expressed in other or more ways than measured (e.g. other disorders) (5) High rates of pre-existing mental health disorders (3) Dose–response phenomena High pre-existing exposure to conflict or other disaster (4) Severe exposure to (potentially) traumatic events during or after the disaster (5) Severe exposure to loss (5) Stressful recovery environment Lack of survival essentials (e.g. clean water, food, health care) (4) Insecurity (5) Severe poverty (4) Low rates of pre-existing mental health disorders (3) Long time since events (opportunity for natural recovery) (1) Dose–response phenomena Limited previous exposure to population-level stressors (2) Limited losses (1) Single event trauma (1) Supportive recovery environment Adequate socio-political response to crisis (3) Security/safety re-established (1) Helpful community self-support mechanisms (3) Availability of effective mental health and social services(3) Helpful cultural or religious coping mechanisms (3) Sampling problems Over-sampling of severely affected areas (1) Over-sampling of likely high risk demographic groups (e.g. women) (1) Socially desirable over-reporting (acquiescence) Expectations raised by the way the overall survey and its questions are framed (3) Interviewer's own expectations (3) Help-seeking behaviour among respondents (3) Measures cover life-time prevalence (1) Measures not fully re-validated for specific emergency-affected context To-be-expected distress (e.g. ongoing experiences of loss and misery) not necessarily consistently distinguished from symptoms of mental disorder (4) Reality-based fears (e.g. due to risk of exposure to ongoing further violence) not necessarily consistently distinguished from symptoms of mental disorder (4) Sampling problems Small sample size Non-probability sampling (3) Low response rates (1) Unreported or low reliability: Internal consistency (1) Test–retest reliability (3) Inter-rater reliability (3) Use of incorrect diagnostic criteria (5) Uncertainty whether translation is adequate (1) Unknown or weak criterion validity of exposure measures (1) Unknown or weak criterion validity of outcome measures (3) Uncertainty whether measures concepts are culturally valid (2) Sampling problems Over sampling of less affected areas (1) Under-sampling of likely high-risk groups (1) Underreporting due to problems in recall (1) Socially under-desirable reporting (social stigma) Declarations of anonymity/confidentiality not persuasive. (3) Lack of privacy in interview settings (3) Measures cover point prevalence (1) Trauma-induced mental disorder expressed in other or more ways than measured (e.g. other disorders) (5) High rates of pre-existing mental health disorders (3) Dose–response phenomena High pre-existing exposure to conflict or other disaster (4) Severe exposure to (potentially) traumatic events during or after the disaster (5) Severe exposure to loss (5) Stressful recovery environment Lack of survival essentials (e.g. clean water, food, health care) (4) Insecurity (5) Severe poverty (4) Low rates of pre-existing mental health disorders (3) Long time since events (opportunity for natural recovery) (1) Dose–response phenomena Limited previous exposure to population-level stressors (2) Limited losses (1) Single event trauma (1) Supportive recovery environment Adequate socio-political response to crisis (3) Security/safety re-established (1) Helpful community self-support mechanisms (3) Availability of effective mental health and social services(3) Helpful cultural or religious coping mechanisms (3) Note: Numbers in brackets reflect a rating of 1–5 of Neugebauer et al.'s2 study, reported in this issue of the Journal (5 = definitive factor in explaining size of obtained rate; 4 = likely factor in explaining size of rate; 3 = possible factor in explaining size; 2 = unlikely factor; 1 = definitively not a factor in explaining size of obtained rate). Schematic overview of possible influences on uncertainty in observed prevalence rates (with ratings applied to the Neugebauer et al.2 study) Sampling problems Over-sampling of severely affected areas (1) Over-sampling of likely high risk demographic groups (e.g. women) (1) Socially desirable over-reporting (acquiescence) Expectations raised by the way the overall survey and its questions are framed (3) Interviewer's own expectations (3) Help-seeking behaviour among respondents (3) Measures cover life-time prevalence (1) Measures not fully re-validated for specific emergency-affected context To-be-expected distress (e.g. ongoing experiences of loss and misery) not necessarily consistently distinguished from symptoms of mental disorder (4) Reality-based fears (e.g. due to risk of exposure to ongoing further violence) not necessarily consistently distinguished from symptoms of mental disorder (4) Sampling problems Small sample size Non-probability sampling (3) Low response rates (1) Unreported or low reliability: Internal consistency (1) Test–retest reliability (3) Inter-rater reliability (3) Use of incorrect diagnostic criteria (5) Uncertainty whether translation is adequate (1) Unknown or weak criterion validity of exposure measures (1) Unknown or weak criterion validity of outcome measures (3) Uncertainty whether measures concepts are culturally valid (2) Sampling problems Over sampling of less affected areas (1) Under-sampling of likely high-risk groups (1) Underreporting due to problems in recall (1) Socially under-desirable reporting (social stigma) Declarations of anonymity/confidentiality not persuasive. (3) Lack of privacy in interview settings (3) Measures cover point prevalence (1) Trauma-induced mental disorder expressed in other or more ways than measured (e.g. other disorders) (5) High rates of pre-existing mental health disorders (3) Dose–response phenomena High pre-existing exposure to conflict or other disaster (4) Severe exposure to (potentially) traumatic events during or after the disaster (5) Severe exposure to loss (5) Stressful recovery environment Lack of survival essentials (e.g. clean water, food, health care) (4) Insecurity (5) Severe poverty (4) Low rates of pre-existing mental health disorders (3) Long time since events (opportunity for natural recovery) (1) Dose–response phenomena Limited previous exposure to population-level stressors (2) Limited losses (1) Single event trauma (1) Supportive recovery environment Adequate socio-political response to crisis (3) Security/safety re-established (1) Helpful community self-support mechanisms (3) Availability of effective mental health and social services(3) Helpful cultural or religious coping mechanisms (3) Sampling problems Over-sampling of severely affected areas (1) Over-sampling of likely high risk demographic groups (e.g. women) (1) Socially desirable over-reporting (acquiescence) Expectations raised by the way the overall survey and its questions are framed (3) Interviewer's own expectations (3) Help-seeking behaviour among respondents (3) Measures cover life-time prevalence (1) Measures not fully re-validated for specific emergency-affected context To-be-expected distress (e.g. ongoing experiences of loss and misery) not necessarily consistently distinguished from symptoms of mental disorder (4) Reality-based fears (e.g. due to risk of exposure to ongoing further violence) not necessarily consistently distinguished from symptoms of mental disorder (4) Sampling problems Small sample size Non-probability sampling (3) Low response rates (1) Unreported or low reliability: Internal consistency (1) Test–retest reliability (3) Inter-rater reliability (3) Use of incorrect diagnostic criteria (5) Uncertainty whether translation is adequate (1) Unknown or weak criterion validity of exposure measures (1) Unknown or weak criterion validity of outcome measures (3) Uncertainty whether measures concepts are culturally valid (2) Sampling problems Over sampling of less affected areas (1) Under-sampling of likely high-risk groups (1) Underreporting due to problems in recall (1) Socially under-desirable reporting (social stigma) Declarations of anonymity/confidentiality not persuasive. (3) Lack of privacy in interview settings (3) Measures cover point prevalence (1) Trauma-induced mental disorder expressed in other or more ways than measured (e.g. other disorders) (5) High rates of pre-existing mental health disorders (3) Dose–response phenomena High pre-existing exposure to conflict or other disaster (4) Severe exposure to (potentially) traumatic events during or after the disaster (5) Severe exposure to loss (5) Stressful recovery environment Lack of survival essentials (e.g. clean water, food, health care) (4) Insecurity (5) Severe poverty (4) Low rates of pre-existing mental health disorders (3) Long time since events (opportunity for natural recovery) (1) Dose–response phenomena Limited previous exposure to population-level stressors (2) Limited losses (1) Single event trauma (1) Supportive recovery environment Adequate socio-political response to crisis (3) Security/safety re-established (1) Helpful community self-support mechanisms (3) Availability of effective mental health and social services(3) Helpful cultural or religious coping mechanisms (3) Note: Numbers in brackets reflect a rating of 1–5 of Neugebauer et al.'s2 study, reported in this issue of the Journal (5 = definitive factor in explaining size of obtained rate; 4 = likely factor in explaining size of rate; 3 = possible factor in explaining size; 2 = unlikely factor; 1 = definitively not a factor in explaining size of obtained rate). We studied the research reported in this issue of the Journal2 and rated different methodological and contextual factors in the study from 1 (not a factor in explaining size of obtained rate) to 5 (definitive factor in explaining size of obtained rate) (see bracketed numbers in Table 1). If no information was available in the paper on an element, then we rated it 3. Starting with the cell in the bottom left, we will discuss here ratings of 4 and 5, which are of main interest in explaining findings. The study2 took place in 1995 in a context (recent genocide that was preceded and followed by violence, fears of revenge killings, ongoing mass displacements, risk of cholera outbreaks, etc.) that not only involved mass loss and trauma but also a highly stressful recovery environment. These risk factors taken together likely explain a substantial part of the rates of mental disorder reported in this study. The cell on the top left draws attention to the thorny issue of distinguishing between (i) mental disorder and (ii) normal (‘understandable’) stress reactions in the face of loss, danger or ongoing stressors—an unresolved issue in psychiatry.3–5 Neugebauer et al. showed some cultural/construct validity of their PTSD measure by demonstrating a predicted dose–response relationship and a predicted gender difference. Yet, it is very plausible that some or many of the symptoms that were reported were not due to exposure to traumatic events during the period of genocide in 1994, but were caused or at least maintained by current stressors in 1995, the year of the interview. For example, a Rwandan woman raped in 1994 may have been interviewed while living in a camp for internally displaced persons in 1995. At the time of the interview, she may have been overwhelmed with fear of going to the bathroom at night, in view of the realistic possibility that rape would re-occur in a highly insecure, poorly lit camp. Similarly, she may have feared that people in the community would find out about the rape and, because of its social stigma, that she would be abandoned by her husband and ostracized by her community—sadly, a plausible scenario.6 Symptoms of current, reality-based (objective) fear—in contrast to symptoms of excessive unrealistic fear—are probably not meant to be interpreted as indicators of mental disorder. In situations of an almost complete breakdown of social structures one may not get meaningful results with self-report symptom checklists. If used, such measures need to be recalibrated by comparing with clinician-administered interviews, such as the Schedules for Clinical Assessment in Neuropsychiatry (SCAN),7 at the time of the survey. We move to the middle column in the table, where we have rated ‘Use of incorrect diagnostic criteria’ as 5, indicating that this is without doubt a methodological factor explaining high rates. Indeed, the authors did not assess PTSD Criterion F (clinical significance) and, ignoring clinical significance as a necessary criterion for DSM-IV diagnosis8 automatically inflates observed DSM-IV rates.9 In short, there were various methodological actors that likely or definitely inflated observed rates. The column on the right describes influences leading to relative lower rates. The one rating of 5 in this column reminds us that the Neugebauer study focuses on only one trauma-induced disorder. Had the study covered more disorders, observed rates of ‘any mental disorder’ would have almost certainly been much higher than the rates reported in this study. In conclusion, we expect that in the sea of Rwandans with genocide-induced symptoms in 1995, there were many genocide-induced emotional disorders, including PTSD, which for a proportion of people would likely have been chronic and severely disabling. Yet, application of the scheme (Table 1) leads us to conclude that we cannot say with much confidence that we have an idea of how many Rwandans had DSM-IV PTSD in 1995. Research efforts in emergencies may perhaps be more productively directed to other questions, such as ‘Are self-report measures valid in these circumstances?’ and ‘What mental health activities are effective in reducing symptoms?’ Of note, whether to focus narrowly on PTSD in emergencies or on a much broader range of mental health and psychosocial problems has been a subject of much debate.10 Most international agencies have recently agreed on a much broader focus.11 Finally, we would like to highlight the enormous value of the researchers’ excellent documentation and analyses of the frequency and nature of traumatic exposures in Rwanda. The authors validated their measurement of traumatic events, which is exceptional. This article includes a unique validated population-based picture of human rights violations during one of the most disturbing pages of human history. We are grateful for comments on a draft of this article by Derrick Silove, Zachary Steel, Peter Ventevogel and Shekhar Saxena. The views expressed in this article are those of the authors solely and do not necessarily represent the views, policies and decisions of the World Health Organization. Conflicts of interest: None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.154 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.004 | 0.007 |
| Open science | 0.009 | 0.003 |
| Research integrity | 0.050 | 0.039 |
| Insufficient payload (model declined to judge) | 0.009 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".