Multi-Call Memory in an AI Care Agent for Chronic Care Management Among Older Adults: Retrospective Observational Study (Preprint)
Bibliographic record
Abstract
Background: Multicall memory capabilities in AI-powered health care communication systems show promise for enhancing patient engagement, but their impact on engagement and patient satisfaction remains unclear. Objective: This study evaluated the relationship between multicall memory usage and key patient experience metrics, including call duration and satisfaction scores, in an AI-powered health care communication system. Methods: We conducted a retrospective analysis of 4415 AI care agent calls from 4189 patients using linear mixed-effects models to account for multiple calls per patient. The primary predictor was the number of memories used per call. Outcomes included call duration (in minutes), net promoter score, and patient satisfaction ratings. We analyzed the full dataset and relevant subsets (completed calls only and memory-using calls only) to assess the robustness of the findings. Results: =.004). Memory usage showed no significant association with patient satisfaction across any analysis. Given that only a small subset of calls used memories and satisfaction data were available only for completed calls, the study may have been underpowered to detect an association between memory use and net promoter score or satisfaction ratings. Conclusions: Multicall memory usage is significantly associated with enhanced behavioral engagement. The findings reveal a disconnect between engagement duration and patient-reported experience, suggesting that memory optimization strategies should focus on behavioral engagement metrics while considering factors beyond usage quantity for patient satisfaction. These results provide evidence-based guidance for health care organizations implementing memory-enabled AI communication systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".