Primary Care EHR data on Social Determinants of Health: Quality and Fitness for Purpose in Precision/Personalised Medicine
Bibliographic record
Abstract
INTRODUCTION: Precision and personalised medicine requires comprehensive genetic, epigenetic, lifestyle, social, community and environmental knowledge of the patient. This approach highlights the importance of the social determinants of health (SDoH), described by the World Health Organization (WHO) as 'the non-medical factors that influence health outcomes, the conditions in which people are born, grow, work, live, and age, and the wider set of forces and systems shaping the conditions of daily life such as economic policies and systems, development agendas, social norms, social policies and political systems'. METHODS: This study examined if countries collect SDoH indicators and, if they do, the quality of the data and whether they are fit for clinical and population health purposes. The sources of data were EHR networks and, where not available, national data collections. RESULTS: While demographic details (age, gender) and rurality were well documented in most countries, we found that data availability and quality for education, occupation, income, socio-economic status, and residential care varied considerably between countries. Data for smoking, obesity, alcohol use, mental health, and substance use were generally poorly recorded. CONCLUSION: Recommendations include a universal set of indicators and taxonomy for SDoH; common data model and metadata standards for national and global harmonisation and monitoring; benchmarks for data quality and fitness-for-purpose; capacity building at national and subnational levels in data collection, data analysis, communication and dissemination of results; ethical and transparent data stewardship; and governance, leadership and diplomacy across multiple sectors to co-create an enabling policy and regulatory environment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".