Assessment of patient handouts on burns created by burn surgeons compared to ChatGPT-4o
Bibliographic record
Abstract
Physicians often face time constraints that may impact the delivery of patient education. Large language models have illustrated promising results in patient education across various specialties. The present study’s aim was to investigate the quality and readability of ChatGPT- generated handouts on burns and compare these results to a published handout. We asked ChatGPT-4o to generate and regenerate patient handouts for seven topics regarding burns. These handouts, along with a patient handout with similar topics published by Hamilton Health Sciences, were assessed. The Quality of Generated Language Outputs for Patients (QGLOP) scale was used to assess handouts based on accuracy/comprehensiveness, bias, currency, and tone, where each domain was scored out of 4 for a total of 16. The Simple Measure of Gobbledygook (SMOG) score was calculated to assess handout readability. The threshold for statistical significance was set at p < 0.05. The mean QGLOP scores for the ChatGPT-4o generated handouts and the published handout did not significantly differ. The mean QGLOP scores between ChatGPT-4o and the published handout were not significantly different for accuracy, bias, currency, and tone. ChatGPT-4o had lower scores on the topic of skin care, but higher scores on coping with burns. The two groups did not significantly differ for any other topic. We found that ChatGPT could produce patient education handouts on burns with scores comparable to those of a patient handout published by a burn unit, suggesting that plastic surgeons would have a similar level of satisfaction for both groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".