Using Systems Engineering to Inform Program Evaluation Practices in Health Professions Education: Conceptualizing Educational Programs as Socio-Technical Systems to Study System Emergence
Bibliographic record
Abstract
Introduction: Understanding the dynamics of educational programs is complex due to the unexpected processes and outcomes resulting from interactions among program elements. In Health Profession Education (HPE), scholars have suggested this complexity might be best studied and understood using program evaluation approaches that capture the program’s planned and emergent processes and outcomes. Doing so might produce a better understanding of how, and what, the program is achieving. However, to date, no such program evaluation framework exists in HPE is able to accomplish this. This dissertation uses System Engineering principles and tools to propose a new program evaluation framework for sensing, characterizing, and explaining the complexity of educational programs. Methods: The Socio-Technical Evaluation of Educational Programs (STEEP) framework was developed and refined by implementing System Engineering principles and tools to study two HPE programs. System emergence was defined as the unintended processes and outcomes in a program. Using a multiple case study approach, these two implementations resulted in a refined STEEP framework including the following methodical steps: relabeling of data, cross-stakeholder analysis, and appraisal of information power related to system emergence. Results: The findings suggest the STEEP framework sensed system emergence in these educational programs, and produced data for characterizing and proposing possible mechanisms to explain the emergence. The results also showed potential sources of system emergence, including convergence and divergence in different stakeholders’ perceptions, as well as the influence of external systems (e.g., a residents’ hospital culture affecting their off-site course experiences). Conclusions: This dissertation positions system emergence as one of the underpinning mechanisms for complex educational programs. Key contributions include a refined STEEP framework, and an improved understanding of mechanisms related to system emergence. These findings provide evaluators and researchers with more refined strategies for studying system emergence. A key recommendation is for program evaluators to continue focusing on identifying interactions between planned and emergent process and outcomes, while also taking the extra analytic step of aiming to clarify the mechanisms driving system emergence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".