Evaluation Models for Undergraduate Nursing Clinical Skills: A Scoping Review Protocol (Preprint)
Bibliographic record
Abstract
BACKGROUND Undergraduate nursing students are expected to perform a high-stakes clinical skills test, which ultimately determines their ability to engage in clinical practice and complete the program. With intake number of students growing nationally, clinical instructors are modifying these skills tests to be shorter in duration as an attempt to meet scheduled class times, severely decreasing the assessments’ accuracy and increasing student stress. OBJECTIVE The objectives of this review are to examine the literature for best practices of evaluation structures or models for undergraduate nursing students participating in clinical skills-based assessments as part of their program curriculum or course evaluation and to explore alternative, best-practice methods for evaluating key clinical skills in nursing students. The goal is to use what is known in the literature to help inform the development of a new evidence-based clinical skills evaluation model that is adaptive to student population growth, enhances the learning experience for students and teaching experience for instructors, and ensures safe patient care. METHODS This scoping review will consider studies involving undergraduate nursing students enrolled in an undergraduate nursing degree program who undergo clinical skills testing as required by the academic institution as part of their degree fulfillment. Original articles published in English from January 1, 2015, through the end of the scoping review period will be included. This review will follow the Joanna Briggs Institute methodology for scoping reviews, as well as the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines. The literature search will utilize the following databases: SCOPUS, CINAHL, Medline, PsycInfo, ProQuest Central, and ProQuest Dissertations and Theses to identify relevant sources. Moreover, non-empirical research, such as editorials, opinion papers, grey literature, and reviews, will be included. Data from professional nursing organizations, including the College of Nurses of Ontario (CNO), the Registered Nurses’ Association of Ontario (RNAO), and the Canadian Association of Schools of Nursing (CASN), will be considered. Two independent reviewers will conduct a screening of article titles and abstracts, followed by full-text reviews. Together, they will determine which articles will proceed to the data extraction stage. If discrepancies arise, a consensus discussion will take place between the two reviewers with assistance from a third reviewer where agreement cannot be reached on article inclusion. RESULTS The project was funded in June 2025. The scoping review is scheduled to begin in July 2025, with the literature search and study selection processes planned through Fall 2025. Results are expected to be submitted for publication in Winter 2026. CONCLUSIONS The results of this review will summarize the current breadth of knowledge on clinical skills testing and best-practice methods used for clinical skill evaluation amongst undergraduate nursing students. Any gaps in the literature will be identified as these can be used to guide future research within this area of nursing education. CLINICALTRIAL The review has been registered in Open Science Framework (OSF). https://doi.org/10.17605/OSF.IO/K7RV4
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.145 | 0.200 |
| Meta-epidemiology (narrow) | 0.004 | 0.004 |
| Meta-epidemiology (broad) | 0.010 | 0.013 |
| Bibliometrics | 0.015 | 0.015 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.007 | 0.006 |
| Insufficient payload (model declined to judge) | 0.054 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".