Team OSCEs: evaluation methodology or educational encounter?
Bibliographic record
Abstract
Context and setting Two university faculties of health sciences piloted the team objective structured clinical examination (TOSCE) format in 2006 to interprofessional undergraduate health sciences students. Most of these students were medical students, but the sample included students drawn from social work, chaplaincy, nursing and occupational therapy (OT) curricula. Subsequently, three TOSCE stations have been delivered over a half-day, conducted every 6 weeks, to medical students at one university since January 2007 as part of the mandatory clerkship curriculum. Other health care students participate as an elective. Both interprofessional groups (medical and nursing students, or social work, OT and chaplaincy students) and uniprofessional groups (medical students, who role-play the roles of other team members) have thus participated in the TOSCE experience at this university for over 16 months. Why the idea was necessary Current interest in interprofessional education for collaborative patient-centred practice (IECPCP) raises challenges of defining the competencies necessary for teamwork and how to teach and then evaluate them. There is little medical education literature that directs curriculum designers to evaluation methods or formative educational assessment tools for IECPCP. What was done Stations for the OSCE were created based on clinical scenarios that require a team approach to care (hence ‘team OSCEs’ [TOSCEs]). Stations are 30 minutes in length and teams of five or six students work through scenarios, depicted by a standardised patient (SP) or a video clip, to determine an interprofessional care plan for the patient involved. Students are evaluated on a set of clinical competencies in an area of focus which varies from station to station (e.g. palliative care), and a set of standardised interprofessional competencies that are consistent across each station. Each station has two evaluators; one is an MD faculty member and one is a faculty member from another allied health profession such as nursing, social work, OT or chaplaincy. Each station includes 10 minutes for feedback given by the SP and the evaluators. Evaluation of results and impact Students and evaluators complete extensive surveys at the end of each TOSCE day. Both students (n = 141) and evaluators (n = 38) have reported a high degree of acceptability of the TOSCE, with 81–100% of respondents responding with ‘agree’ or ‘strongly agree’ to a series of acceptability questions. Similarly, the majority of both student and evaluator respondents (79–100%) agreed or strongly agreed that the TOSCE format was quite feasible. Of note, the students felt the 10 minutes of feedback following each station was amongst the most useful learning they had received in their training. Three more TOSCE stations are being introduced to evaluate the reliability and validity of the TOSCE. Student TOSCE scores are being compared with those on multiple-choice question tests and clinical application exercises in the same content areas. Thirty students will complete six TOSCE stations as part of this evaluation, which will include randomisation so that the effect of the group versus the individual can be investigated. The TOSCE holds promise for learners at all levels for a variety of clinical scenarios where both health care content and team-based skills are necessary. At the least, it is a formative educational tool, and current reliability and validity data will determine its effectiveness as an evaluation methodology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.072 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".