Characteristics of Surgical Coaching Interventions: A Systematic Review
Bibliographic record
Abstract
OBJECTIVE: Coaching is increasingly utilized as an educational intervention for performance improvement in surgeons and surgical trainees. Surgical coaching has been utilized across a broad range of specialties, experience levels and outcomes with generally positive results. Coaching interventions are often developed by individual institutions for their own context which has resulted in a heterogenous group of interventions. This review aims to investigate surgical coaching interventions to identify common characteristics that comprise an effective coaching intervention. METHODS: A systematic review was conducted to identify studies investigating surgical coaching interventions up to July 2024. Studies were limited to English language peer-reviewed studies that adequately described the characteristics and outcomes of the surgical coaching intervention. Data on the primary and secondary outcomes, study objective and participants' demographics were also recorded. RESULTS: The search across 4 electronic databases generated 9538 citations. Following screening and review of full text articles 28 studies were included in the review. Surgical coaching interventions were carried out in 8 separate countries with the majority (22/28) in North America. Studies involved between 3 and 107 participants. Coaching interventions were markedly heterogenous, and specific details of the methods used were inconsistently documented. Study length ranged from 1 (9/28) to 14 (1/28) sessions and duration from less than 15 minutes (1/29) to greater than 3 hours (3/28). The most common themes were goal setting (10/28), feedback (7/28) and reflection (7/28). Outcomes were generally positive with 47 of 55 identified outcomes demonstrating benefit from surgical coaching. There were 3 key domains and 13 sub-domains that comprised the majority of coaching interventions. CONCLUSIONS: Surgical coaching has been shown to be a promising intervention that requires more rigorous research to develop the field. We have identified 3 key domains which can be utilized to analyses and develop coaching interventions in the future.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".