Assessing Patient-Reported Satisfaction With Care and Documentation Time in Primary Care Through AI-Driven Automatic Clinical Note Generation: Protocol for a Proof-of-Concept Study
Bibliographic record
Abstract
BACKGROUND: Relisten is an artificial intelligence (AI)-based software developed by Recog Analytics that improves patient care by facilitating more natural interactions between health care professionals and patients. This tool extracts relevant information from recorded conversations, structuring it in the medical record, and sending it to the Health Information System after the professional's approval. This approach allows professionals to focus on the patient without the need to perform clinical documentation tasks. OBJECTIVE: This study aims to evaluate patient-reported satisfaction and perceived quality of care, assess health care professionals' satisfaction with the care provided, and measure the time spent on entering records into the electronic medical record using this AI-powered solution. METHODS: This proof-of-concept (PoC) study is conducted as a multicenter trial with the participation of several health care professionals (nurses and physicians) in primary care centers (CAPs). The key outcome measures include (1) patient-reported quality of care (evaluated through anonymous surveys), (2) health care professionals' satisfaction with the care provided (assessed through surveys and structured interviews), and (3) time saved on clinical documentation (determined by comparing the time spent manually writing notes versus reviewing and correcting AI-generated notes). Statistical analyses will be performed for each objective, using independent sample comparison tests according to normality evaluated with the Kolmogorov-Smirnov test and Lilliefors correction. Stratified statistical tests will also be performed to consider the variance between professionals. RESULTS: The protocol has been developed using the SPIRIT (Standard Protocol Items: Recommendations for Interventional Trials) checklist. Recruitment began in July 2024, and as of November 2024, a total of 318 patients have been enrolled. Recruitment is expected to be completed by March 2025. Data analysis will take place between April and May 2025, with results expected to be published in June 2025. CONCLUSIONS: We expect an improvement in the perceived quality of care reported by patients and a significant reduction in the time spent taking clinical notes, with a saving of at least 30 seconds per visit. Although a high quality of the notes generated is expected, it is uncertain whether a significant improvement over the control group, which is already expected to have high-quality notes, will be demonstrated. TRIAL REGISTRATION: ClinicalTrials.gov NCT06618092; https://clinicaltrials.gov/study/NCT06618092. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/66232.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.045 | 0.028 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.004 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.027 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".