Framework to prioritize health outcomes of particulate matter exposure using national claims data
Bibliographic record
Abstract
OBJECTIVES: Although particulate matter (PM) exposure poses significant public health risks, previous research has focused on limited clinical areas. However, emerging evidence and pathological mechanisms of PM suggest that PM may exert broader systemic effects across a wide range of diseases. Therefore, we aim to identify and prioritize research questions to evaluate health impacts of PM exposure across various clinical specialties. METHODS: A structured collaborative process was conducted between April and November 2024 in South Korea, incorporating systematic literature reviews, multidisciplinary expert discussions, and knowledge-sharing seminars. The primary outcomes were the identification of diseases potentially influenced by PM exposure and the development of corresponding research questions. The literature review synthesized more than 417 publications, including the U.S. Environmental Protection Agency's integrated science assessment materials, a government-issued abstract compendium on PM covering 2010-2019, and studies published from 2020 to 2024 identified via a structured search. These were categorized by exposure duration (short- or long-term) and diseases outcome (incidence or progression). Prioritization was based on three criteria: pathological causality, clinical impact (public health burden), and feasibility using the Korea National Health Insurance Service (K-NHIS). RESULTS: A total of 99 experts from epidemiology, data science, and 14 clinical specialties participated. The experts panel (mean age: 46.1 years; mean professional experience: 20.5 years) identified 211 research questions across 80 diseases. These were classified by disease outcome: disease incidence (short-term, 54; long-term, 64) and progression (short-term, 47; long-term, 46). Notably, several clinical areas such as ophthalmology, dermatology, and otolaryngology were underrepresented. CONCLUSION: This structured, multidisciplinary approach broadened the scope of PM-related clinical research beyond commonly studied clinical area. This scalable framework can be adapted in other regions with similar claims data systems to guide evidence-based research agendas and inform public health policies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.057 | 0.065 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.007 |
| Bibliometrics | 0.027 | 0.011 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.005 | 0.007 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".