Contemporary Methods for Assessment of Undergraduate Medical Students
Notice bibliographique
Résumé
INTRODUCTION The multiple purposes of assessment of undergraduate medical students can be clustered under three categories or approaches: Assessment for learning (formative/diagnostic assessment); Assessment as learning (reflective/self-assessment); and Assessment of learning (summative assessment) (1). Assessment for learning (formative/diagnostic assessment) is useful both for lecturers to modify their teaching methods to be more effective for educators and for students to adapt their learning strategies to be more efficient learners. Using proper search design and tools, teachers can identify what students know, when and how they learn best, and assess their ability to apply knowledge in real-life. They can also find knowledge gaps, misconceptions, and confusion. Thus, teachers can improve teaching by reorganizing approaches and resources and giving feedback to help students enhance their learning strategies. Assessment for learning occurs throughout teaching and can include focused questioning, discussions, complex tasks, and various other activities, such as essays, portfolios, videos, role-play, and debates, in addition to formal exams like mini-clinical evaluation exercises (min-CEX). These assessments should be followed by individualized feedback. In assessment as learning (reflective/self-assessment), students act as their own assessors, reflecting and analyzing their learning to make improvements. This involves identifying strengths and areas for growth through formative and summative assessments, adjusting strategies rooted in metacognitive monitoring. Students need training to reflect on their academic and non-academic performance to provide effective responses. Key attributes include self-awareness and critical analysis, involving documenting incidents, reflecting, analyzing for improvements, and acting on insights. Assessment of learning (summative assessment) provides evidence of the achievements of students against standards and outcomes. It is summative, conducted at set points during or at the end of training, to decide students’ achievement levels, progress, or course completion. It typically includes written and clinical exams, such as multiple-choice, essay questions, objective structured clinical examination (OSCE), mini-CEX, and case exams. Interpreting assessment data requires careful analysis before making critical decisions. Results are examined separately, and students often must perform satisfactorily in certain areas to progress. In competency-based education, students must demonstrate adequate achievement in specific skills to pass. Programmatic assessment (PA) effectively collates and interprets data from various assessment methods during a course, providing feedback and supporting credible decisions (2). The student’s performance from a single assessment (e.g., OSCE) is called a ‘data point’(3). Each data point is used in the assessment for learning, thereby guiding students to improve their learning (2). The accumulation of multiple data points over a period of time is scrutinized by competence committees to make high-stake decisions, such as a pass or fail status (4). It is emphasized that the pass or fail decisions must be made by a group of assessment experts and not by individual assessors (5). Realizing that each single assessment method has its own limitations and if used alone to make high-stake decisions will compromise the reliability of the assessment process, PA uses a ‘menu of assessment methods’ to cover the shortcomings of individual assessment tools (5). PA requires advanced planning, thoughtful selection of assessment methods, careful scheduling of assessment plans and feedback approaches (5). Therefore, implementing this requires several modifications to traditional assessment methods and the allocation of additional resources. This resource demand can be particularly challenging for many medical schools, especially in developing countries. Mukurunge et al (6) have aptly explained it as follows: “There should be a well-established support structure for educators, a supportive administrative department in the institution, and an established group of experts who will make high-stakes assessment decisions affecting students’ progression in the programme. Additional aspects that should be in place include training workshops for educators, mentors, and preceptors in the clinical area, as well as timely and constructive feedback after each assessment and the use of multiple methods of assessment for the collection of data on student performance” (6). The usage of multiple methods of assessment or data points results in a large volume of information on student performance. This in turn, would necessitate to establish an information management system to collate and analyse the data (7). The success of PA implementation is not universal and poses a challenge for resource-limited institutions. However, thoughtful planning and careful scheduling, along with a data information management system, may make it possible to implement it successfully. REFERENCES Earl L, Katz, S. (2006) Rethinking Classroom Assessment with Purpose in Mind: Assessment for Learning, Assessment as Learning, Assessment of Learning. Canada: Manitoba Education, Citizenship and Youth, Winnipeg. 2006. Torre D, Rice NE, Ryan A, Bok H, Dawson LJ, Bierer B, et al. Ottawa 2020 consensus statements for programmatic assessment – 2. Implementation and practice. Med Teach. 2021;43(10):1149–60. Shrivastava SR, Shrivastava PS. Programmatic assessment of medical students: pros and cons. J Prim Health Care: Open Access. 2018;8(3):1–2. Govaerts M, Van der Vleuten C, Schut S. Implementation of programmatic assessment: challenges and lessons learned. Educ Sci. 2022;12(717):1–6. Heeneman S, de Jong LH, Dawson LJ, Wilkinson TJ, Ryan A, Tait GR, et al. Ottawa 2020 consensus statement for programmatic assessment – 1. Agreement on the principles. Med Teach. 2021;43(10):1139–48. Mukurunge E, Nyoni CN, Hugo L. Assessment approaches in undergraduate health professions education: towards the development of feasible assessment approaches for low-resource settings. BMC Medical Education. 2024;24:318 https://doi.org/10.1186/s12909-024-05264-x Ryan A, Terry J. From traditional to programmatic assessment in three (not so) easy steps. Educ Sci. 2022;12(487):1–13.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,016 | 0,049 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,010 | 0,007 |
| Études des sciences et des technologies | 0,002 | 0,002 |
| Communication savante | 0,005 | 0,003 |
| Science ouverte | 0,004 | 0,005 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,047 | 0,019 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».