From Patient Voices to Policy: Data Analytics Reveals Patterns in Ontario’s Hospital Feedback
Bibliographic record
Abstract
Abstract Patient satisfaction is a central measure of high-performing healthcare systems, yet real-world evaluations at scale remain challenging. In this study, we analyzed over 120,000 de-identified patient reviews from 45 Ontario hospitals between 2015 and 2022. We applied natural language processing (NLP), including named entity recognition (NER), to extract insights on hospital wards, patient health outcomes, and medical conditions. We also examined regional demographic data to identify potential disparities emerging during the COVID-19 pandemic. Our findings show that nearly 80% of the hospitals studied had fewer than 50% positive reviews, exposing systemic gaps in meeting patient needs. In particular, negative reviews decreased during COVID-19, suggesting possible shifts in patient expectations or increased appreciation for strained healthcare workers; however, certain units, such as intensive care and cardiology, experienced fewer positive ratings, reflecting pandemic and related pressures on critical care services. ‘Anxiety’ emerged as a recurrent concern in negative reviews, pointing to the growing awareness of mental health needs. Furthermore, hospitals located in regions with higher percentages of visible minority and low-income populations initially saw higher positive review rates before COVID-19, but this trend reversed after 2020. Collectively, these results demonstrate how large-scale unstructured data can identify fundamental drivers of patient satisfaction, while underscoring the urgent need for adaptive strategies to address anxiety and combat systemic inequalities. Author Summary Understanding what patients think and feel about hospital care can lead to better health services and outcomes. We analyzed more than 120,000 patient reviews from 45 Ontario hospitals between 2015 and 2022. Our study combined natural language processing techniques to identify key concerns, including anxiety, billing difficulties, and interactions with staff. We also compared patient experiences before and during the COVID-19 pandemic, uncovering a drop in negative reviews and a rise in positive reviews, though certain units—such as intensive care—faced growing pressure. A particularly revealing finding was that hospitals located in regions with higher numbers of visible minority and low-income groups received more positive feedback before the pandemic, but this reversed after 2020. These patterns hint at deeper systemic issues, especially during times of crisis. By pinpointing the main drivers of satisfaction and dissatisfaction, our work highlights the need for healthcare services that prioritize kindness, clear communication, efficient operations, and equitable access for all. Lessons from this research could guide targeted improvements, ensuring that every patient, regardless of background or income, receives the compassionate and timely care they deserve. Our hope is that policymakers, hospital administrators, and community advocates will use these findings to shape policies that improve patient trust and well-being.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.075 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".