Patient-Generated Collections for Organizing Electronic Health Record Data to Elevate Personal Meaning, Improve Actionability, and Support Patient–Health Care Provider Communication: Think-Aloud Evaluation Study
Bibliographic record
Abstract
BACKGROUND: Through third party applications, patients in the United States have access to their electronic health record (EHR) data from multiple health care providers. However, these applications offer only a predefined organization of these records by type, time stamp, or provider, leaving out meaningful connections between them. This prevents patients from efficiently reviewing, exploring, and making sense of their EHR data based on current or ongoing health issues. The lack of personalized organization and important connections can limit patients' ability to use their data and make informed health decisions. OBJECTIVE: To address these challenges, we created Discovery, an experimental app that enables patients to organize their medical records into collections, analogous to placing pictures in photo albums. These collections are based on the evolving understanding of the patients' past and ongoing health issues. The app also allows patients to add text notes to collections and their constituent records. By observing how patients used features to select records and assemble them into collections, our goal was to learn about their preferred mechanisms to complete these tasks and the challenges they would face in the wild. We also intended to become more informed about the various ways in which patients could and would like to use collections. METHODS: We conducted a think-aloud evaluation study with 14 participants on synthetic data. In session 1, we obtained feedback on the mechanics for creating and assembling collections and adding notes. In session 2, we focused on reviewing collections, finding data patterns within them, and retaining insights, as well as exploring use cases. We conducted reflexive thematic analysis on the transcribed feedback. RESULTS: Collections were useful for personal use (quick access to information, reflection on medical history, tracking health, journaling, and learning from past experiences) and clinical visits (preparation and raising physicians' awareness). Assembling EHR data into reliable collections could be difficult for typical patients due to considerable manual work and lack of medical knowledge. However, automated collection building could alleviate this issue. Furthermore, having EHR data organized in collections may have limited use. However, augmenting them with patient-generated data, which are entered with flexible richness and structure, could add context, elevate meaning, and improve actionability. Finally, collections might produce a misconstrued health picture, but bringing the physician in the loop for verification could increase their clinical validity. CONCLUSIONS: Collections can be a powerful tool for advancing patients' proactivity, awareness, and self-advocacy, potentially facilitating patient-centered care. However, patients need better support for incorporating their own everyday data and adding meaningful annotations for future reference. Improvements in the comprehensiveness, efficiency, and reliability of the collection assembly process through automation are also necessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.041 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".