Linking Electronic Health Records and In-Depth Interviews to Inform Efforts to Integrate Social Determinants of Health Into Health Care Delivery: Protocol for a Qualitative Research Study
Bibliographic record
Abstract
BACKGROUND: Health systems are attempting to capture social determinants of health (SDoH) in electronic health records (EHR) and use these data to adjust care plans. To date, however, methods for identifying social needs, which are the SDoH prioritized by patients, have been underexplored, and there is little guidance as to how clinicians should act on SDoH data when caring for patients. Moreover, the unintended consequences of collecting and responding to SDoH are poorly understood. OBJECTIVE: The objective of this study is to use two data sources, EHR data and patient interviews, to describe divergences between the EHR and patient experiences that could help identify gaps in the documentation of SDoH in the EHR; highlight potential missed opportunities for addressing social needs, and identify unintended consequences of efforts to integrate SDoH into clinical care. METHODS: We are conducting a qualitative study that merges discrete and free-text data from EHRs with in-depth interviews with women residing in rural, socioeconomically deprived communities in the Mid-Atlantic region of the United States. Participants had to confirm that they had at least one visit with the large health system that serves the region. Interviews with the women included questions regarding health, interaction with the health system, and social needs. Next, with consent, we extracted discrete data (eg, diagnoses and medication orders) for each participant and free-text clinician notes from this health system's EHRs between 1996 and the year of the interview. We used a standardized protocol to create an EHR narrative, a free-text summary of the EHR data. We used NVivo to identify themes in the interviews and the EHR narratives. RESULTS: To date, we have interviewed 88 women, including 51 White women, 19 Black women, 14 Latina women, 2 mixed Black and Latina women, and 2 Asian Pacific women. We have completed the EHR narratives on 66 women. The women range in age from 18 to 90 years. We found corresponding EHR data on all but 4 of the interview participants. Participants had contact with a wide range of clinical departments (eg, psychiatry, neurology, and infectious disease) and received care in various clinical settings (eg, primary care clinics, emergency departments, and inpatient hospitalizations). A preliminary review of the EHR narratives revealed that the clinician notes were a source of data on a range of SDoH but did not always reflect the social needs that participants described in the interviews. CONCLUSIONS: This study will provide unique insight into the demands and consequences of integrating SDoH into clinical care. This work comes at a pivotal point in time, as health systems, payors, and policymakers accelerate attempts to deliver care within the context of social needs. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/36201.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.100 | 0.072 |
| Meta-epidemiology (narrow) | 0.003 | 0.004 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.005 | 0.006 |
| Science and technology studies | 0.009 | 0.005 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.005 | 0.008 |
| Insufficient payload (model declined to judge) | 0.045 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".