The Prevalence of Rheumatoid Arthritis: A Systematic Review of Population-based Studies
Bibliographic record
Abstract
OBJECTIVE: To estimate the prevalence of rheumatoid arthritis (RA) from international population-based studies and investigate the influence of prevalence definition, data sources, classification criteria, and geographical area on RA prevalence. METHODS: A search of ProQuest, MEDLINE, Web of Science, and EMBASE was undertaken to identify population-based studies investigating RA prevalence between 1980 and 2019. Studies were reviewed using the Joanna Briggs Institute approach for the systematic review and Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. RESULTS: Sixty studies met the inclusion criteria. There was a wide range of point prevalence reported (0.00-2.70%) with a mean of 0.56% (SD 0.51) between 1986 and 2014, and a mean period prevalence of 0.51% (SD 0.35) between 1955 and 2015. RA point and period prevalence was higher in urban settings (0.69% vs 0.48%) than in rural settings (0.54% vs 0.25%). An RA diagnosis validated by rheumatologists yielded the highest period prevalence of RA and was observed in linked databases (0.80%, SD 0.1). CONCLUSION: The literature reports a wide range of point and period prevalence based on population and method of data collection, but average point and period prevalence of RA were 51 in 10,000 and 56 in 10,000, respectively. Higher urban vs rural prevalence may be biased due to poor case findings in areas with less healthcare or differences in risk environment. The population database studies were more consistent than sampling studies, and linked databases in different continents appeared to provide a consistent estimate of RA period prevalence, confirming the high value of rheumatologist diagnosis as classification criteria.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.086 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.012 | 0.012 |
| Bibliometrics | 0.021 | 0.020 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".