Developing a Multimodal Screening Algorithm for Mild Cognitive Impairment and Early Dementia in Home Health Care: Protocol for a Cross-Sectional Case-Control Study Using Speech Analysis, Large Language Models, and Electronic Health Records
Bibliographic record
Abstract
BACKGROUND: Mild cognitive impairment and early dementia (MCI-ED) are frequently unrecognized in routine care, particularly in home health care (HHC), where clinical decisions are made under time constraints and cognitive status may be incompletely documented. Federally mandated HHC assessments, such as the Outcome and Assessment Information Set (OASIS), capture health and functional status but may miss subtle early cognitive changes. Speech, language, and interactional patterns during routine patient-nurse communication, together with information embedded in unstructured clinical notes, may provide complementary signals for earlier identification. OBJECTIVE: This protocol describes the development and evaluation of a multimodal screening approach for identifying MCI-ED in HHC by integrating (1) speech and interaction features from routine patient-nurse encounters (verbal communication), (2) large language model-based extraction of MCI-ED-related information from HHC notes and encounter transcripts, and (3) structured variables from OASIS. METHODS: This ongoing cross-sectional case-control study is being conducted in collaboration with VNS Health (formerly Visiting Nurse Service of New York). Eligible participants are adults aged ≥60 years receiving HHC services. Case/control assignment uses a 2-stage process: electronic health record (EHR) prescreening followed by clinician-reviewed cognitive assessment (Montreal Cognitive Assessment and Clinical Dementia Rating) for consented participants without an existing mild cognitive impairment diagnosis. For Aim 1, each participant contributes 3 audio-recorded routine patient-nurse encounters linked to EHR data, including OASIS and free-text clinical notes. Aim 1 extracts acoustic, linguistic, emotional, and interactional features from patient-nurse verbal communication. Aim 2 uses a schema-guided large language model pipeline to extract and normalize MCI-ED-related symptoms, lifestyle risk factors, and communication deficits from HHC notes and encounter transcripts, supported by a human-annotated gold-standard dataset. Aim 3 integrates speech, extracted text variables, and OASIS predictors using supervised machine learning with stratified nested cross-validation; evaluation will include discrimination, calibration, and subgroup performance checks across race, sex, and age. RESULTS: Between February 2024 and July 2025, a total of 114 HHC patients completed study-administered cognitive assessments and were classified as 55 MCI-ED cases and 59 cognitively normal controls. Audio-recorded patient-nurse encounters had a median duration of 19 (IQR 12-23) minutes and a median of 56 (IQR 31-80) utterances per encounter; nurses contributed more words than patients (median 842, IQR 461-1218 vs median 589, IQR 303-960). In exploratory feasibility analyses, multimodal models integrating speech, interactional features, and structured EHR/OASIS variables outperformed single-source models. CONCLUSIONS: This protocol describes a reproducible multimodal framework for MCI-ED screening in HHC using routinely generated data streams. Initial implementation results support feasibility of data collection and end-to-end processing and suggest potential value of integrating interactional speech features with clinical text and OASIS variables. Final model evaluation, subgroup analyses, and validation will follow the prespecified analytic procedures on the finalized study dataset. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/82731.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.067 | 0.063 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.023 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".