Using natural language processing to compare task‐specific verbal cues in coached versus noncoached cardiac arrest teams during simulated pediatrics resuscitation
Bibliographic record
Abstract
OBJECTIVES: Coaches improve cardiopulmonary (CPR) outcomes in real-world and simulated settings. To explore verbal feedback that targets CPR quality, we used natural language processing (NLP) methodologies on transcripts from a published pediatric randomized trial (coach vs. no coach in simulated CPR). Study objectives included determining any differences by trial arm in (1) overall communication and (2) metrics over minutes of CPR and (3) exploring overall frequencies and temporal patterns according to degrees of CPR excellence. METHODS: A human-generated transcription service produced 40 team transcripts. Automated text search with manual review assigned semantic category; word count; and presence of verbal cues for general CPR, compression depth or rate, or positive feedback to transcript utterances. Resulting cue counts per minute (CPM) were corresponded to CPR quality based on compression rate and depth per minute. CPMs were compared across trial arms and over the 18 min of CPR. Adaptation to excellence was analyzed across four patterns of CPR excellence determined by k-shape methods. RESULTS: Overall coached teams experienced more rate-directive, depth-directive, and positive verbal cues compared with noncoached teams. The frequency of coaches' depth cues changed over minutes of CPR, indicating adaptation. In coached teams, the number of depth-directive cues differed among the four patterns of CPR excellence. Noncoached teams experienced fewer utterances by type, with no adaptation over time or to CPR performance. CONCLUSION: NLP extracted verbal metrics and their patterns in resuscitation sessions provides insight into communication patterns and skills used by CPR coaches and other team members. This could help to further optimize CPR training, feedback, excellence, and outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".