Google star ratings of Canadian hospitals: a nationwide cross-sectional analysis
Bibliographic record
Abstract
BACKGROUND: Data on patients' self-reported hospital experience can help guide quality improvement. Traditional patient survey programmes are resource intensive, and results are not always publicly accessible. Unsolicited online hospital reviews are an alternative data source; however, the nature of online reviews for Canadian hospitals is unknown. METHODS: We conducted a nationwide cross-sectional study of Canadian acute care hospitals with more than 10 Google Reviews during the 2018-2019 fiscal year. We characterised the volume and distribution of Google Reviews of Canadian hospitals, and assessed their correlation with hospital characteristics (teaching status, size, occupancy rate, length of stay, resource utilisation) and Canadian Patient Experience Survey on Inpatient Care (CPES-IC) scores. RESULTS: 167 out of 523 (31.9%) acute care hospitals in Canada met the inclusion criteria. Among included hospitals, there was a total of 10 395 Google Reviews and a median of 35 reviews per hospital. The mean Google Star Rating for included hospitals was 2.85 out of 5, with a range of 1.36-4.57. Teaching hospitals had significantly higher mean Google Star Ratings compared with non-teaching hospitals (3.16 vs 2.81, p <0.01). There was a weak, positive correlation between hospitals' Google Star Ratings and CPES-IC 'Overall Hospital Experience' scores (p =0.04), but no significant correlation between Google Star Ratings and other hospital characteristics or subcategories of CPES-IC scores. INTERPRETATION: There is significant interhospital variation in patients' self-reported care experiences at Canadian acute care hospitals. Online reviews can serve as a readily accessible source of real-time data for hospitals to monitor and improve the patient experience.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".