Global Disparities in Simulation-Based Learning Performance: Serial Cross-Sectional Mixed Methods Study
Bibliographic record
Abstract
Abstract Background Simulated programs provide health care professionals (HCPs) with a learning opportunity to develop clinical competencies and improve patient outcomes in a safe and controlled environment. While the benefits of simulation training are well established, there is a paucity of research assessing its differential impact, if any. SIMBA (Simulation via Instant Messaging for Bedside Application) provides simulation-based learning through WhatsApp and Zoom (Zoom Video Communications, Inc) to increase HCPs’ confidence in managing various medical conditions. Objectives This study aims to explore whether there are differences in the clinical performance of HCPs participating in SIMBA sessions based on gender, country of work, and training grade. Methods This study assessed participants in 17 SIMBA sessions from May 2020 to June 2022. WhatsApp chats containing participants’ approach to the simulated scenarios were graded using an adapted version of the Global Rating Scale consisting of 6 domains: eliciting history; physical examination; investigations, diagnostic tests, and imaging; interpretation of investigations and imaging; clinical judgment; and management and follow-up or discharge plan. These domains were rated using a Likert-type scale of 1 (not done) to 5 (excellent) prior to the session based on expert inputs. All WhatsApp transcripts were evaluated against the scale postsimulation session. Unadjusted and adjusted means and 95% CIs of the scores for the 6 performance variables were calculated using multiple linear regression models. The P value for heterogeneity between the mean performance scores was calculated using likelihood ratio tests by using an analysis of variance. Results A total of 289 participants across 49 countries who completed pre-SIMBA and post-SIMBA surveys in the 17 simulation sessions were included in the analysis. Participants from high-income countries scored higher in all categories of the Global Rating Scale (GRS) except the physical examination and interpretation score. Junior-grade participants scored significantly higher in history taking (junior=4.2, middle=3.7, and senior=3.7; P =.003) and physical examination (junior=4.0, middle=3.7, and senior=3.5; P =.068), but this was not significantly different. There were no statistically significant differences in GRS scores between male and female participants. Conclusions The significant differences in clinical performance scores between low- and middle-income countries and high-income countries highlight the need for better medical education resources to bridge existing gaps in health care globally. The decrease in some clinical competency scores following career progression could be addressed by simulation-based training to maintain the same quality of history taking and physical examination skills. These outcomes, including no gendered differences in simulation-based learning, hold profound implications for tailoring medical education strategies, fostering equitable training, and elevating patient care standards on a global scale. The need for targeted interventions and capacity-building efforts via context-specific training and tailored approaches to health care education is emphasized.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".