Health Reporting from Different Data Sources: Does it Matter for Mental Health?
Bibliographic record
Abstract
BACKGROUND: Mental disorders are typically stigmatized conditions associated with negative stereotypes, which may lead individuals to underreport them. Thus, survey data may be subject to biases. Although administrative data has some limitations, it is an alternative data source that may be considered more objective. AIMS OF THE STUDY: This study aimed to identify the degree of agreement between survey and administrative health care data for mental health conditions, factors affecting underreporting, and whether underreporting also occurs for physical health conditions. METHODS: We used Ontario data from the Canadian Community Health Survey linked to health records to examine the presence of mental health conditions (i.e., schizophrenia and mood disorders) and select physical health conditions (i.e., diabetes and cancer). Using administrative data as the reference standard, we created four categories for each health condition based on the level of agreement between the two data sources: consistent cases and non-cases (i.e. individuals with concordant data based on their reported health condition), and people who were found to underreport and overreport a condition (i.e. where the condition was present in the administrative data, but not in the survey data and vice-versa, respectively). The overall level of agreement was assessed using Cohen's kappa statistic. Probit regressions were estimated to determine the factors affecting underreporting. RESULTS: The Kappa statistics for mood disorder was fair (k= 0.26) and moderate for schizophrenia (k = 0.49). Physical health conditions had higher kappa values (diabetes, k = 0.81; ever having cancer, k = 0.68), with the exception of currently having cancer (k = 0.24). Underreporting was highest for the most stigmatizing condition, schizophrenia (63%), followed by mood disorders (39%) and cancer (39%), and lowest for diabetes (25%). Older age, being born in Africa and Asia, and being employed all increased the probability of underreporting among individuals identified in the administrative data; the opposite held for social assistance. DISCUSSION: We extended previous work on mental health reporting by combining survey data with administrative data to examine the level of agreement between respondents' self-reported mental health and administrative records. The data include some mental disorders not studied previously. We examined the entire adult population; this is important because prevalence of schizophrenia may be less common among older population groups due to higher mortality among this patient population. Additionally, there may be potential age-related differences in stigma and mental health conditions. The administrative health data captured only health services covered by the public provincial health insurance plan and thus did not capture medical care provided by psychologists, social workers, and nurses. While this would affect Kappa statistic values, it does not directly affect the underreporting analyses. IMPLICATIONS FOR HEALTH CARE PROVISION AND USE: Our results suggest that disclosure of mental health conditions may differ by the level of stigma, which has implications for obtaining accurate estimates of mental health prevalence from self-reported data sources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.433 | 0.701 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.007 | 0.019 |
| Science and technology studies | 0.003 | 0.007 |
| Scholarly communication | 0.009 | 0.010 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".