Web search trends on fibromyalgia: development of a machine learning model
Bibliographic record
Abstract
OBJECTIVES: Fibromyalgia (FM) is a chronic pain condition characterised by widespread musculoskeletal pain, fatigue, and cognitive dysfunction. The growing reliance on the internet for health-related information has transformed how individuals seek medical knowledge, particularly for complex conditions like FM. This study aimed to analyse online search behaviours related to FM across multiple countries, identify temporal trends, and assess machine learning models for predicting search interest. METHODS: Google Trends data (2020-2024) were analysed across sixteen countries. Time-series analysis, linear regression, and the Mann-Kendall trend test assessed monotonic trends, while seasonal decomposition identified periodic fluctuations. An Auto-Regressive Integrated Moving Average (ARIMA) model forecasted search volumes for 2025. Machine learning models, including Random Forest (RF) and Extreme Gradient Boosting (XGBoost), were used to predict search trends, with feature importance evaluated using SHAP (Shapley Additive Explanations) values. RESULTS: Search interest in FM varied across countries, with China, the UK, the USA and Canada showing the highest engagement, while Peru, Spain and Turkey had the lowest. Brazil, Italy and the UK exhibited rising search trends, whereas Argentina, Canada, Greece and the USA showed declines. Seasonal analysis revealed mid-year peaks in Brazil and Italy, while Turkey saw late autumn increases. ARIMA forecasting predicted stable or increasing trends in Brazil, Canada and Mexico, while Germany and Venezuela showed slight declines. Machine learning analysis identified short-term search history (search volumes from the previous day, week, and month) as the most influential predictor. CONCLUSIONS: Understanding online search behaviour can enhance FM education. Targeted awareness campaigns and improved digital health literacy initiatives could sustain engagement and improve patient knowledge. Future efforts should focus on optimising online health resources and integrating evidence-based decision aids.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".