Evaluation of Incremental Validity of Casper in Predicting Program and National Licensure Performance of Undergraduate Nursing Students: Protocol for a Mixed Methods Study
Bibliographic record
Abstract
BACKGROUND: Academic success has been the primary criterion for admission to many nursing programs. However, academic success as an admission criterion may have limited predictive value for success in noncognitive skills. Adding situational judgment tests, such as Casper, to admissions procedures may be one strategy to strengthen decisions and address the limited predictive value of academic admission criteria. In 2021, admissions processes were modified to include Casper based on concerns identified with noncognitive skills. OBJECTIVE: This study aims to (1) assess the incremental validity of Casper scores in predicting nursing student performance at years 1, 2, 3, and 4 and on the National Council Licensing Examination (NCLEX) performance; and (2) examine faculty members' perceptions of student performance and influences related to communication, professionalism, empathy, and problem-solving. METHODS: We will use a multistage evaluation mixed methods design with 5 phases. At the end of each year, students will complete questionnaires related to empathy and professionalism and have their performance assessed for communication and problem-solving in psychomotor laboratory sessions. The final phase will assess graduate performance on the NCLEX. Each phase also includes qualitative data collection (ie, focus groups with faculty members). The goal of the focus groups is to help explain the quantitative findings (explanatory phase) as well as inform data collection (eg, focus group questions) in the subsequent phase (exploratory sequence). All students enrolled in the first year of the nursing program in 2021 were asked to participate (n=290). Faculty will be asked to participate in the focus groups at the end of each year of the program. Hierarchical multiple regression will be conducted for each outcome of interest (eg, communication, professionalism, empathy, and problem-solving) to determine the extent to which scores on Casper with admission grades, compared to admission grades alone, predict nursing student performance at years 1-4 of the program and success on the national exam. Thematic analysis of focus group transcripts will be conducted using interpretive description. The quantitative and qualitative data will be integrated after each phase is complete and at the end of the study. RESULTS: This study was funded in September 2021, and data collection began in March 2022. Year 1 data collection and analysis are complete. Year 2 data collection is complete, and data analysis is in progress. CONCLUSIONS: At the end of the study, we will provide the results of a comprehensive analysis to determine the extent to which the addition of scores on Casper compared to admission grades alone predicts nursing student performance at years 1-4 of the program and on the NCLEX exam. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): RR1-10.2196/48672.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.106 | 0.095 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.005 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.005 | 0.003 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.026 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".