Variability in Global Prevalence of Interstitial Lung Disease
Bibliographic record
Abstract
There are limited epidemiologic studies describing the global burden and geographic heterogeneity of interstitial lung disease (ILD) subtypes. We found that among seventeen methodologically heterogenous studies that examined the incidence, prevalence and relative frequencies of ILDs, the incidence of ILD ranged from 1 to 31.5 per 100,000 person-years and prevalence ranged from 6.3 to 71 per 100,000 people. In North America and Europe, idiopathic pulmonary fibrosis and sarcoidosis were the most prevalent ILDs while the relative frequency of hypersensitivity pneumonitis was higher in Asia, particularly in India (10.7-47.3%) and Pakistan (12.6%). The relative frequency of connective tissue disease ILD demonstrated the greatest geographic variability, ranging from 7.5% of cases in Belgium to 33.3% of cases in Canada and 34.8% of cases in Saudi Arabia. These differences may represent true differences based on underlying characteristics of the source populations or methodological differences in disease classification and patient recruitment (registry vs. population-based cohorts). There are three areas where we feel addition work is needed to better understand the global burden of ILD. First, a standard ontology with diagnostic confidence thresholds for comparative epidemiology studies of ILD is needed. Second, more globally representative data should be published in English language journals as current literature has largely focused on Europe and North America with little data from South America, Africa and Asia. Third, the inclusion of community-based cohorts that leverage the strength of large databases can help better estimate population burden of disease. These large, community-based longitudinal cohorts would also allow for tracking of global trends and be a valuable resource for collective study. We believe the ILD research community should organize to define a shared ontology for disease classification and commit to conducting global claims and electronic health record based epidemiologic studies in a standardized fashion. Aggregating and sharing this type of data would provide a unique opportunity for international collaboration as our understanding of ILD continues to grow and evolve. Better understanding the geographic and temporal patterns of disease prevalence and identifying clusters of ILD subtypes will facilitate improved understanding of emerging risk factors and help identify targets for future intervention.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".