Tracking Population-Level Anxiety Using Search Engine Data: Ecological Study
Bibliographic record
Abstract
BACKGROUND: Anxiety disorders are the most prevalent mental disorders globally, with a substantial impact on quality of life. The prevalence of anxiety disorders has increased substantially following the COVID-19 pandemic, and it is likely to be further affected by a global economic recession. Understanding anxiety themes and how they change over time and across countries is crucial for preventive and treatment strategies. OBJECTIVE: The aim of this study was to track the trends in anxiety themes between 2004 and 2020 in the 50 most populous countries with high volumes of internet search data. This study extends previous research by using a novel search-based methodology and including a longer time span and more countries at different income levels. METHODS: We used a crowdsourced questionnaire, alongside Bing search query data and Google Trends search volume data, to identify themes associated with anxiety disorders across 50 countries from 2004 to 2020. We analyzed themes and their mutual interactions and investigated the associations between countries' socioeconomic attributes and anxiety themes using time-series linear models. This study was approved by the Microsoft Research Institutional Review Board. RESULTS: Query volume for anxiety themes was highly stable in countries from 2004 to 2019 (Spearman r=0.89) and moderately correlated with geography (r=0.49 in 2019). Anxiety themes were predominantly long-term and personal, with "having kids," "pregnancy," and "job" the most voluminous themes in most countries and years. In 2020, "COVID-19" became a dominant theme in 27 countries. Countries with a constant volume of anxiety themes over time had lower fragile state indexes (P=.007) and higher individualism (P=.003). An increase in the volume of the most searched anxiety themes was associated with a reduction in the volume of the remaining themes in 13 countries and an increase in 17 countries, and these 30 countries had a lower prevalence of mental disorders (P<.001) than the countries where no correlations were found. CONCLUSIONS: Internet search data could be a potential source for predicting the country-level prevalence of anxiety disorders, especially in understudied populations or when an in-person survey is not viable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".