How Words Matter - Coaches’ Verbal Interventions and Their Relationship to Coach/Client Outcomes
Bibliographic record
Abstract
While coaching has been widely adopted by organizations as a human resource development strategy, in-session coach interventions remain opaque. This study examined links between specific coach verbal behaviors and independent session ratings from clients and coaches. Forty-eight coach-client pairs contributed recordings and ratings for three sessions (2-4 of a six-session series). Transcripts were coded with a 37-category scheme. For each code we derived two metrics—absolute frequency (interventions/hour) and within-session percentage of coach talk—and used median splits to compare high versus low groups on three outcomes: goal alignment, task accomplishment, and perceived bond. Mann-Whitney U tests supported several patterns. Client ratings were higher when empathic, meaning-making behaviors—reflections of feeling, explorations of emotion and learning, affirming feedback, and judicious humor—comprised a larger share of coach talk, and lower when advice, disconfirming feedback, ecosystem probing, repeated verification/checking, or tangential thought occurred at higher rates. Coaches reported greater alignment (and, by percentage, task accomplishment) with more minimal encouragers, and higher task accomplishment and bonding with more affirmations, humor, and learning exploration; they reported lower alignment when emotions or values were explored more frequently. Behavioral dysregulation was associated with lower coach-rated task accomplishment. Longer sessions (~45 vs. ~30 minutes) were associated with higher client ratings on all three outcomes and with higher coach-rated task accomplishment. Overall, findings suggest that style composition and dosage matter, and that client and coach perspectives may be complementary rather than interchangeable. Coaching practice may benefit from prioritizing empathic reflections and focused explorations while limiting high-dosage influence moves.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".