Internet search results correlate with seasonal variation of sarcoidosis
Bibliographic record
Abstract
BACKGROUND: The etiology and pathophysiology of sarcoidosis remains unclear, with epidemiologic studies limited by its relatively low prevalence. The internet has prompted patients to seek information about medical diagnoses online; Google Trends provides access to an anonymized version of this data, which has a new role in epidemiology. We hypothesize that there is seasonal variation in the relative search interest of sarcoidosis, which would suggest seasonal variation in the incidence of sarcoidosis. METHODS: Google Trends was used to assess the relative search volume from 2010 to 2020 for "sarcoidosis" and "sarcoid" in 7 countries. ANOVA with multiple comparisons was performed to compare the mean relative search volume by month and by season for each country, with a p-value less than 0.05 indicating statistical significance. RESULTS: Our analysis revealed a significant seasonal variation in search popularity in 4 of the 7 countries and in the Northern Hemispheric countries combined. Direct comparison showed search terms to be more popular in spring, specifically March & April, than in the winter. Southern Hemisphere data was not statistically significant but showed a trend towards a nadir in December and a peak in September and October. CONCLUSIONS: Overall, these findings suggest seasonal variation with a possible peak in spring and nadir in winter. This supports the hypothesis that sarcoidosis has seasonal variation and is more commonly diagnosed in spring, but more evidence is needed to support this, as well as investigation into the pathophysiology of sarcoidosis to explain this phenomenon.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".