Contemporary Methods for Assessment of Undergraduate Medical Students
Bibliographic record
Abstract
INTRODUCTION The multiple purposes of assessment of undergraduate medical students can be clustered under three categories or approaches: Assessment for learning (formative/diagnostic assessment); Assessment as learning (reflective/self-assessment); and Assessment of learning (summative assessment) (1). Assessment for learning (formative/diagnostic assessment) is useful both for lecturers to modify their teaching methods to be more effective for educators and for students to adapt their learning strategies to be more efficient learners. Using proper search design and tools, teachers can identify what students know, when and how they learn best, and assess their ability to apply knowledge in real-life. They can also find knowledge gaps, misconceptions, and confusion. Thus, teachers can improve teaching by reorganizing approaches and resources and giving feedback to help students enhance their learning strategies. Assessment for learning occurs throughout teaching and can include focused questioning, discussions, complex tasks, and various other activities, such as essays, portfolios, videos, role-play, and debates, in addition to formal exams like mini-clinical evaluation exercises (min-CEX). These assessments should be followed by individualized feedback. In assessment as learning (reflective/self-assessment), students act as their own assessors, reflecting and analyzing their learning to make improvements. This involves identifying strengths and areas for growth through formative and summative assessments, adjusting strategies rooted in metacognitive monitoring. Students need training to reflect on their academic and non-academic performance to provide effective responses. Key attributes include self-awareness and critical analysis, involving documenting incidents, reflecting, analyzing for improvements, and acting on insights. Assessment of learning (summative assessment) provides evidence of the achievements of students against standards and outcomes. It is summative, conducted at set points during or at the end of training, to decide students’ achievement levels, progress, or course completion. It typically includes written and clinical exams, such as multiple-choice, essay questions, objective structured clinical examination (OSCE), mini-CEX, and case exams. Interpreting assessment data requires careful analysis before making critical decisions. Results are examined separately, and students often must perform satisfactorily in certain areas to progress. In competency-based education, students must demonstrate adequate achievement in specific skills to pass. Programmatic assessment (PA) effectively collates and interprets data from various assessment methods during a course, providing feedback and supporting credible decisions (2). The student’s performance from a single assessment (e.g., OSCE) is called a ‘data point’(3). Each data point is used in the assessment for learning, thereby guiding students to improve their learning (2). The accumulation of multiple data points over a period of time is scrutinized by competence committees to make high-stake decisions, such as a pass or fail status (4). It is emphasized that the pass or fail decisions must be made by a group of assessment experts and not by individual assessors (5). Realizing that each single assessment method has its own limitations and if used alone to make high-stake decisions will compromise the reliability of the assessment process, PA uses a ‘menu of assessment methods’ to cover the shortcomings of individual assessment tools (5). PA requires advanced planning, thoughtful selection of assessment methods, careful scheduling of assessment plans and feedback approaches (5). Therefore, implementing this requires several modifications to traditional assessment methods and the allocation of additional resources. This resource demand can be particularly challenging for many medical schools, especially in developing countries. Mukurunge et al (6) have aptly explained it as follows: “There should be a well-established support structure for educators, a supportive administrative department in the institution, and an established group of experts who will make high-stakes assessment decisions affecting students’ progression in the programme. Additional aspects that should be in place include training workshops for educators, mentors, and preceptors in the clinical area, as well as timely and constructive feedback after each assessment and the use of multiple methods of assessment for the collection of data on student performance” (6). The usage of multiple methods of assessment or data points results in a large volume of information on student performance. This in turn, would necessitate to establish an information management system to collate and analyse the data (7). The success of PA implementation is not universal and poses a challenge for resource-limited institutions. However, thoughtful planning and careful scheduling, along with a data information management system, may make it possible to implement it successfully. REFERENCES Earl L, Katz, S. (2006) Rethinking Classroom Assessment with Purpose in Mind: Assessment for Learning, Assessment as Learning, Assessment of Learning. Canada: Manitoba Education, Citizenship and Youth, Winnipeg. 2006. Torre D, Rice NE, Ryan A, Bok H, Dawson LJ, Bierer B, et al. Ottawa 2020 consensus statements for programmatic assessment – 2. Implementation and practice. Med Teach. 2021;43(10):1149–60. Shrivastava SR, Shrivastava PS. Programmatic assessment of medical students: pros and cons. J Prim Health Care: Open Access. 2018;8(3):1–2. Govaerts M, Van der Vleuten C, Schut S. Implementation of programmatic assessment: challenges and lessons learned. Educ Sci. 2022;12(717):1–6. Heeneman S, de Jong LH, Dawson LJ, Wilkinson TJ, Ryan A, Tait GR, et al. Ottawa 2020 consensus statement for programmatic assessment – 1. Agreement on the principles. Med Teach. 2021;43(10):1139–48. Mukurunge E, Nyoni CN, Hugo L. Assessment approaches in undergraduate health professions education: towards the development of feasible assessment approaches for low-resource settings. BMC Medical Education. 2024;24:318 https://doi.org/10.1186/s12909-024-05264-x Ryan A, Terry J. From traditional to programmatic assessment in three (not so) easy steps. Educ Sci. 2022;12(487):1–13.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.049 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.010 | 0.007 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.004 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.047 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".