Data-Driven Activities Involving Electronic Health Records: An Activity and Task Analysis Framework for Interactive Visualization Tools
Bibliographic record
Abstract
Electronic health records (EHRs) can be used to make critical decisions, to study the effects of treatments, and to detect hidden patterns in patient histories. In this paper, we present a framework to identify and analyze EHR-data-driven tasks and activities in the context of interactive visualization tools (IVTs)—that is, all the activities, sub-activities, tasks, and sub-tasks that are and can be supported by EHR-based IVTs. A systematic literature survey was conducted to collect the research papers that describe the design, implementation, and/or evaluation of EHR-based IVTs that support clinical decision-making. Databases included PubMed, the ACM Digital Library, the IEEE Library, and Google Scholar. These sources were supplemented by gray literature searching and reference list reviews. Of the 946 initially identified articles, the survey analyzes 19 IVTs described in 24 articles that met the final selection criteria. The survey includes an overview of the goal of each IVT, a brief description of its visualization, and an analysis of how sub-activities, tasks, and sub-tasks blend and combine to accomplish the tool’s main higher-level activities of interpreting, predicting, and monitoring. Our proposed framework shows the gaps in support of higher-level activities supported by existing IVTs. It appears that almost all existing IVTs focus on the activity of interpreting, while only a few of them support predicting and monitoring—this despite the importance of these activities in assisting users in finding patients that are at high risk and tracking patients’ status after treatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.026 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.005 |
| Bibliometrics | 0.017 | 0.008 |
| Science and technology studies | 0.002 | 0.006 |
| Scholarly communication | 0.010 | 0.009 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".