The Impact of an Alternative Approach to Computing Station Cut Scores in an OSCE
Bibliographic record
Abstract
The OSCE is gaining widespread recognition as a valid means of assessing entry-to-practice competence, or eligibility for licensure, of physicians, physiotherapists, and other health professionals. Given the high-stakes nature of these licensure OSCEs, robust psychometric properties of the exams are essential. One of these properties is the resistance of cut scores used in determining pass—fail decisions to such sources of error as differences in examiner perceptions of competence and examiner stringency in judging competence. A number of standard-setting methods have been described in the literature on performance-based assessment. Methods are typically categorized as relative or absolute,1,2 with most administrators responsible for high-stakes examinations preferring absolute or criterion-referenced methods. Absolute standard-setting methods compare candidates' performances with an externally determined or defined measure (criterion) and are typically categorized as test-centered or examinee-centered.1,2 These categories distinguish methods according to whether judgments about competence are based primarily on inspection of test items (e.g., Angoff, Ebel, and similar methods) or on judgments about examinees (e.g., contrasting groups, borderline group). The common elements for all methods include (1) use of expert judges and (2) reference to a hypothetical “minimally competent” person or a hypothetical “borderline competent” performance.1 Descriptions and classifications of standard-setting methods can be found in review articles by Cizek,1 Berk,3 and Cusimano.4 A modification of the mean-borderline-group method that is now being employed by a number of credentialling agencies entails identification of a subgroup of candidates actually performing the exam who are identified by the examiners as having a level of clinical competence that is just on the borderline between being competent and not being competent. The station scores for this borderline group are averaged to generate the station cut score. In this approach, the rating of candidates' performances as competent, borderline, or not competent is concurrent with completion of checklists and/or other scoring rubrics by these same examiners. This modified mean-borderline-group method has potentially interesting implications for the determination of cut scores when large-scale, multi-site examinations are employed. The cut score is calculated as the mean of the scores of all candidates who receive borderline ratings, regardless of site of administration (or examiner). Thus, in a multi-site examination where examiners are nested within sites, an examiner who identifies a greater number of “borderline competent” candidates during the exam has a greater influence than other examiners on the resultant cut score for that station. The inequality of examiners' influence over station cut scores is inconsistent with other standard-setting methods described, such as the Angoff and Ebel methods, and could be problematic when combined with the potential for examiners' differing perceptions of what constitutes a borderline performance. The proposed alternative to this current approach is one in which every examiner's opinion or concept of borderline competence is weighted the same. In other words, the mean of each examiner's borderline group is calculated first, then the mean across examiners evaluating the same station is calculated to determine the cut score. The impacts of individual examiners on the resultant cut score are thus equalized. The effect of the alternative method on cut scores and the practical impact on pass-fail decisions was explored in order to determine whether further investigation of cut score validity is required. Method Data for 1,373 candidates who participated in four administrations (years) of an OSCE used in a national physiotherapy examination were used in the study. Each administration of the OSCE consisted of 20 stations in which candidates were required to perform a clinical skill in the context of a clinical scenario. Results from two stations had been removed for administrative reasons, so these results were not included in the data set provided. Therefore, the scores for a total of 78 stations were used in the study. Candidates rotated through circuits of ten ten-minute stations and ten five-minute stations. Each site consisted of two ten-minute circuits per five-minute circuit, or two examiners of each ten-minute station, and one examiner of each five-minute station. Each year had a median of 40 candidates per site. (Therefore, as a rule, ten-minute station examiners evaluated 20 candidates and five-minute station examiners evaluated 40 candidates.) There were seven, five, 12, and 14 sites of administration in the four respective years of the exam. Each candidate was allowed to choose the sites at which he or she participated, so assignment of candidates to sites was not a random process and will likely have been influenced by location of training. The examiners were clinicians from the local community. They were assigned to stations according to their self-identified areas of clinical expertise. They attended a training session where they oriented to the examination procedures and scoring processes before the exam. Scoring of the OSCE. The candidates performances were rated using a task-specific dichotomous checklist where clinician examiners record whether key behaviors are demonstrated correctly.* In addition, overall performance was rated on a six-point rating scale. The two middle anchors (3 of 6 and 4 of 6) on this scale were “borderline unsatisfactory” and “borderline satisfactory.” The borderline group used in computation of cut scores is considered to include all candidates assigned either of these two scores. The over-all rating of performance was considered for the sole purpose of identifying borderline candidates to calculate cut scores. Computation of Cut Scores. The traditional approach entailed finding the mean checklist score for all candidates identified as borderline, regardless of site or examiner. The alternative approach entailed finding the examiner-specific mean checklist score for the borderline group, then computing the average of these means across examiners. This second method, in effect, weights all examiners' opinions equally. These two checklist-based cut scores were computed for the 78 stations used over the four examinations. In addition to the overall examination score, candidates are required to perform satisfactorily in a criterion number of stations to pass. The number of stations required to pass an examination fluctuates from year to year, depending on the level of difficulty of the examination. In the hypothetical situation constructed for the purpose of these analyses, I used 12- and 13-station criteria to examine two different scenarios for the impact of using the alternative method of computation on exam-level pass—fail decisions. Results When I examined descriptive data for examiner patterns with regard to use of the “borderline competent” rating, the findings were very consistent for the five-minute and ten-minute stations. For all analyses presented, the results are pooled across the 78 stations to maximize power. There were observed differences in the proportions of candidates deemed borderline by different examiners examining the same station. The mean discrepancy (the range in the proportions of candidates identified as borderline by different examiners of a given station) across the 78 stations was 48%, with the lowest discrepancy being 11% (where an examiner at one site identified no candidate as borderline, and an examiner at another site identified 11% as borderline) and the highest discrepancy being 90% (with one examiner identifying no borderline candidate and another identifying 90% of candidates as borderline). Clearly, in the computation of cut scores, some examiners are contributing substantially more borderline candidates than others. The examiner-specific cut scores (the mean checklist score for borderline candidates at a site) also displayed within-station ranges. For example, there was a cut-score discrepancy of 56% (of total possible checklist points) between two examiners of a given station on two of the 78 stations examined. Thirty-five of the 78 stations (45%) had cut score ranges of 30% or more across examiners. Thus, when discrepancies in examiner-specific cut scores were considered in conjunction with discrepancies in the proportions of candidates rated by individual examiners as borderline there was high potential that, in this hypothetical case based on checklist scores only, use of the proposed alternative method of computation could lead to very different results, at least at the level of station cut scores and pass—fail decisions. The remaining analyses assessed the impact of the observed variability among examiners in their applications of the borderline rating on exam results. Impact on Raw Cut Scores. The two computation methods generated very similar ranges of station cut scores across the 78 stations. The traditional method resulted in cut scores ranging from 37.03% for the most difficult station to 87.13% for the easiest, while the alternative method resulted in cut scores ranging from 35.37% to 86.93%. At the level of station, differences between the two cut scores were strikingly small, with the maximum difference in cut scores between the two methods being 4.73%. Differences between cut scores were relatively normally distributed around a mean difference of 0.31. There was no significant difference between the cut-scores using the two different approaches (paired t77 = 1.66, p =.10). Also worth noting is the fact that of the 78 stations, 22 were used on more than one occasion. From these stations, it was possible to assess the stability of cut scores over time. The absolute difference in cut scores for the two iterations of each station was calculated for each method. The two methods showed equal levels of cut-score stability of the 22 stations over multiple occasions, in that there was no significant difference in the absolute difference score (reflective of cut score change over time) between methods (paired t21 = 1.01, p =.33). Impact on Station-level Pass—Fail Decisions. There was an equally small effect on pass—fail decisions made at the level of station. For 58 of the 78 stations (74%) there was perfect concordance of pass—fail decisions made using the two methods (i.e., failure rates were unaffected by use of the alternative method). For the 20 stations that were affected, the alternative method increased failure rates at 13 stations while decreasing rates at seven stations. Changes in failure rates for the 20 affected stations ranged from 2% to 18% of candidates examined. There was, however, no significant difference in station failure rates between methods (paired t77 = 0.78, p =.44). When the 22 repeat-use stations were examined for stability of station-level pass—fail decisions, the alternative computation method did not result in a substantial practical effect. The absolute differences between failure rates at two occasions of station use were no different when compared for the two methods (paired t21 = 0.24, p =.82). Impact on Exam-level Pass—Fail Decisions. The examination-level pass—fail decisions are made on the basis of the number of stations passed. In other words, their station scores must be above the station cut score on, say, 12 of 20 stations. In the hypothetical situation created for the study, candidates were required to meet a criterion of passing 12 or 13 stations (both scenarios were examined), based on historical precedent at this and other similar testing organizations. When candidate performance data were used to examine whether the effect of using the alternative cut-score computation method on station-level decisions would translate into an effect at the level of the entire examination, very little impact was seen. The exam-level agreement rates for the two methods (i.e., the proportion of candidates where the exam-level pass—fail decision was unaffected) ranged from 95% to 98% for the four exams, with kappa coefficients ranging from 0.88 to 0.96. Pooled across all four years of administration, the agreement rates were 96% for a 13-station criterion and 97% for a 12-station criterion. Discussion There are a number of standard-setting methods described in the OSCE and performance-based—assessment literature. The currently used modification of the mean-borderline-group method of computing station cut scores in multi-site examinations is the only method that gives unequal weighting to judges (examiners) as a result of the practicalities of implementation. The observed ranges of examiner-specific cut scores, combined with the differences in proportions of borderline candidates identified by different examiners of the same station, open up the potential for individual examiner(s) to influence the cut score in a manner inconsistent with the opinions of the other examiners. This is a sharp contrast with most other standard-setting models, where all experts' judgments are weighted equally. This study proposed an alternative method that attempts to correct for this imbalance and examined the practical implications of using this alternative. Given the observed variability of examiners in the application of the borderline rating, there was surprisingly little impact when empirical data were subjected to the alternative computation method. There was remarkable consistency between the cut scores generated by the two methods within stations, with the largest observed difference being only 5%, or the equivalent of no more than two checklist items. The small differences translated to similarly consistent pass—fail decisions at the level of individual stations. Of the 78 stations on which the two methods were compared, decisions were unaffected at 58 (74%). Furthermore, neither raw cut scores nor station-level failure rates were systematically affected by use of the alternative method. That is, equal weighting of examiner opinions did not consistently result in more or less stringent cut scores. The already small effect was further attenuated at the level of pass—fail decisions for the four 19- or 20-station examinations. Concordance rates between the two methods were very high, with final decisions being unaffected for 97% of all candidates included in the analyses. It appears that because there is no systematic effect of the alternative method at the station level, increases in failure rates at one station are being counteracted by decreases at another station in the same examination. Further, equal weighting of examiner opinions does not influence the reproducibility of station cut scores (and the resultant failure rates) across testing occasions. The degree to which failure rates changed from one use of a given station to the next was not systematically altered by introduction of the new computation method. It should be noted that this study examined only the effect of the alternative computation method on checklist-based ratings of candidate performance. Station composite scores, which the literature suggests are more reliable,5 often form the basis for both station-level and exam-level decisions. The station composite scores were not examined in this study due to the complicating effects of measuring multiple constructs with multiple scoring rubrics. The extent to which the findings on checklist scores would be replicated on station composites is still untested but may be of interest to test developers and administrative bodies that use composite scores to assess performances on multidimensional examinations. The extent to which the observed differences in borderline ratings are related to differences in the candidates' abilities across sites or differences among examiners in their use of the “borderline competent” rating may be of interest to some but is essentially an academic argument. Reasons for observed variations in the frequencies of borderline ratings used by examiners have not been studied to date, and could not be determined through the study design used in this project. Variability in applying the borderline rating was in fact observed in the data set used for the study. For these data, despite differences among examiners, the practical implications of weighting their opinions according to liberality of use of the “borderline” rating do not suggest a need to change current practices in large-scale, multi-site OSCEs. Given the equivocal psychometric benefits of one approach versus the other, a decision about which computation method should be employed when using the mean-borderline-group technique should be based on philosophical and practical rationales.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.195 | 0.441 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.011 | 0.012 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.011 | 0.005 |
| Open science | 0.007 | 0.007 |
| Research integrity | 0.004 | 0.007 |
| Insufficient payload (model declined to judge) | 0.016 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".