Standardised versus individualised assessment: related problems divided by a common language
Bibliographic record
Abstract
This article is part of a series in Medical Education entitled ‘Dialogue’. Each publication in the series will be a transcription of an e-mail discussion about a current issue in the field held by two scholars who have approached the issue from different perspectives. For further details, see the editorial published in Med Educ 2012;46(9):826–7. In this volume, Lambert Schuwirth, Professor of Medical Education at Flinders University in Adelaide and David Swanson, Vice President of Assessment Programs on the National Board of Medical Examiners Philadelphia, discuss the tensions between adopting assessment protocols that are tailored to the needs and abilities of the individual learner and those that are structured and standardised in a manner that makes them more generically applicable to all learners within a particular population. I really liked your presentation at the Ottawa conference in Kuala Lumpur, where you demonstrated that one cannot simply run an analysis without being sufficiently clear about the underlying assumptions, or as you stated it: ‘You have to tell the analysis what to analyse.’ It made me think about some of the assumptions underlying our theory and practice of assessment. One of the most insightful pieces I have read this year was Kane’s chapter on validity in Educational Measurement.1 It made a couple of things clear to me. First, why universe representation, or the inference from observed scores (the scores based on the particular observations/items that were in the actual examination; the sample) to universe scores (the scores that would be obtained had the examination contained all possible observations/items; the ‘population’ scores), is such a crucial aspect of the validity argument. It helped me to understand why we always claim that a test cannot be valid if it is not reliable: namely, because one of the central inferences is flawed, the whole argument cannot be made. However, a test can be reliable but not valid if one of the other inferences is flawed but the inference of universe representation is not. It also made clear how reliability – as one approach to universe representation – fits into this train of argument. Yet there is a contradiction for me that I have been unable to resolve. In our thinking about the universe of admissible observations/items (not all items/observations are admissible, only those that are relevant for what we want to assess; much like in epidemiology the ‘population at risk’), we seem to widely acknowledge that it is diverse. Isn’t this the reason why we blueprint to make sure that we cover various aspects of the universe? So we assume that we achieve a better universe representation by adding heterogeneity to our sample. Yet in classical test theory and in generalisability theory, we seem to seek for high inter-item correlations, internal consistency and low within-subject inter-item variance. This strikes me as being based on the assumption of the universe being homogeneous. So I am puzzled: do our scientific theories start from the notion of a heterogeneous or a homogeneous universe? Can two such basic assumptions – peacefully – co-exist? What do you think? Thanks for the kind words. I tried to illustrate that although it is very common for those analysing the reliability (generalisability) of an objective structured clinical examination (OSCE) to do the analysis as though everyone took the examination at the same time with the same cases, standardised patients (SPs) and examiners at each station, the reality is that this is almost never true. Especially in large-scale OSCEs, multiple circuits (or diets) involving different SPs and examiners are necessary to test everyone within a reasonable timeframe, even if the same cases are used throughout. If calculated incorrectly, reliability estimates are likely to suggest the examination is much more precise than is actually the case. I’m also a big fan of Mike Kane’s chapter on validity in the latest tome of Educational Measurement.1 It is a very readable treatment of a complex subject, providing an excellent (and brief) historical perspective on how the concepts of validity and validation have evolved over the last century. It also provides a useful way to think about the components of a good validity argument, which those engaging in validation research can readily translate into practical research efforts. Kane’s concept of threats (challenges) to validity1 is particularly helpful; in a sense, the use of multiple circuits, SPs and examiners in an OSCE can be viewed as threatening the validity of the interpretation of OSCE scores, which leads to research efforts to quantify the magnitude of the associated sources of measurement error. With regard to the contradiction you described, I also find this troubling, although I tend to think of it as a tension, rather than a contradiction, between what you would like to measure (i.e. the inferences you would like to draw from scores) and what you sometimes have to be satisfied with due to real-world constraints on testing time and resources. I think that Kane’s presentation of the components of the ‘interpretive argument’1 is very helpful in this regard. As you said, a richer universe representation is achieved by making test content more heterogeneous, and this tends to strengthen the extrapolation step in the interpretive argument because the included ‘test tasks’ are broader and, often, more like the real world. This has consequences for the generalisation step in the argument, both because richer test tasks typically take more testing time (lessening the amount of measurement information yielded per unit of time) and because increasing heterogeneity tends to result in lower inter-task correlations. Test blueprints provide a disciplined way to define the ‘sampling plan’ for a test, and then generalisability theory can be used to estimate the strength of relationships between (randomly parallel) replications in the form of generalisability coefficients and standard errors of measurement. I think G theory provides the necessary conceptual tools to look at more and less heterogeneous conceptions of what is to be measured and the resulting impact on the reproducibility of test scores. But the narrowing of test content and restricting of the methods used often come at a price, weakening the extrapolation argument. The designer of an assessment has to decide where the ‘sweet spot’ is – the point at which what’s measured is rich enough to be of interest, yet reproducible enough to serve the purposes for which scores will be used. As an example, a test designer might be interested in measuring students’ skills in interpreting the results of diagnostic studies. The initial blueprint may allow for the inclusion of a very broad, diverse range of diagnostic studies, and the reproducibility of test scores may be low because individual students are good at interpreting some types of study but not others. Restricting the blueprint to include only chest X-rays, electrocardiograms (ECGs), electrolytes and arterial blood gas (ABG) findings in specified proportions may substantially improve the reproducibility if each of these can individually be assessed reproducibly in limited testing time. This is because replications of the assessment procedure will have the same structure, and better reproducibility will be observed even if correlations between skills in the interpretation of chest X-rays, ECGs, electrolytes and ABG findings are quite low. But many would view the inferences that can be drawn from the resulting scores as inconsistent with the original intent of the assessment, not extrapolating to the broader domain actually of interest. Thank you very much for your answer; it makes things much clearer to me. After all, I am not a psychometrician but an md, and I may just be showing my limited understanding of psychometrics here. If I understand you correctly, it is actually the same as with research: the more you focus your research on a local situation, the richer – and probably more ecologically valid – the information, but the lower the generalisability to other situations. So it is a matter of trying to find the sweet spot between being optimally rich in information and optimally generalisable. But there is still something nagging in my brain: namely, that reproducibility of a sample result would be the only approach to universe representation. And this has to do with my background as an md. In this line of work, clinicians are always dealing with individual patients and the goal is to help each one of them achieve an optimal health situation. So clinicians are always in a sort of n = 1 situation. For this, they need to collect information, some of which may be ‘graded on the curve’ and compared with results in a reference population, such as laboratory tests. Some information, on the other hand, will be based on individual judgements, such as pathology reports. If the data represent the former, clinicians want them to be objective – they do not want the laboratory analyst’s opinion – and they want them to be numerical. If the latter, clinicians DO want the opinion and not a number. It is only in these forms, one more ‘objective’ and one more ‘subjective’, that data can have meaning to clinicians and allow them to form an optimal universe representation (meaning the most complete picture of the single disease or combination of diseases the patient has). In addition, clinicians are rarely tempted to arithmetically combine lab values just because they are of the same format (with the exception perhaps of the anion gap); they almost always combine information from different sources and modalities. This may seem difficult, but I assure you that clinicians have no problem with combining a glucose level of 35 mmol/L with complaints of thirst and fatigue and the finding of poorly healing wounds in conjunction with absent peripheral arterial pulsations into diabetes mellitus. In fact, they have to do this to be able to use all the different modalities to optimise the diagnostic procedure and treatment of this n = 1 patient. Now here is my problem. As an educator, I actually have to deal with individual students; regardless of the numbers that enrol into university, my duty is to ensure that each and every one of them reaches his or her potential. If I were just to combine test results because they are derived from the same format (multiple-choice questions with multiple-choice questions, OSCE stations with OSCE stations), I would be ignoring the quite convincing literature stating that the content, not the format, determines what a test measures, and therefore combining content-similar pieces of information is more sensible than combining format-similar pieces.2,3 Also, I would miss out on the opportunity to use this collected information to optimise and individualise assessment for each and every student. So, this is where my concern comes from: isn’t the tension we have been talking about just a result of our approach to assessment as a testing approach instead of a diagnostic approach because it is the testing approach that forces us to combine things that do not belong together? This would be analogous to developments in our research field, where it is increasingly acknowledged that one-off big-bang studies are not sufficient to provide meaningful answers and that, especially for more complicated concepts, a programme of research is indispensable. This would mean for me that a complicated concept such as ‘competence’ requires a programme of assessment in which content-similar rather than just format-similar elements are combined. I know I may be pushy, but this leads to one other issue: namely, the notion of equity. I think all assessment developers agree that the assessment must be fair and equitable. But if I understand it correctly, equitable does not only mean treating equal people equally, but also treating unequal people unequally. I remember a joke about a very traditional men’s choir that was forced to accept women and reluctantly agreed. Its admission criteria, however, remained the same: any successful applicant should be able to sing a bass, baritone or tenor piece from sheet music prima vista. If this makes you smile, you agree with me that standardisation is not the only route to equity. The other part of the equation involves catering to differences between people, without treating these differences automatically as deficiencies, and retaining the quality of all judgements. This, again, relates to clinical patient care: diagnosing and treating every patient precisely according to the same algorithm would NOT do justice to individual differences; it is actually in the adaptation of the expert to each individual patient that quality and equity in health care are achieved. This is why I like to say that a checklist in assessment is never a good substitute for assessor expertise. I’m sorry this was a little long and perhaps a bit preachy, but I think it also demonstrates the differences in the remits of our work and why they are so complementary. Let me start by saying that I don’t think of myself as a psychometrician either – and my psychometric colleagues at the National Board of Medical Examiners (NBME) agree: they have often advised me to stop practising psychometrics without a license. As an md with an excellent background in psychometrics, I think your understanding of the practical measurement issues important in medical education clearly surpasses my own. I think it is very useful to think about the characteristics and uses of educational assessments as analogous to the use of diagnostic studies in making clinical decisions, although the sensitivity, specificity and predictive value of the latter are typically much, much better. Performance on an educational assessment seems to me to be like the result of a (single) diagnostic test. And, like skilled clinicians, teachers have access to a rich array of diverse information for developing an understanding of students’ strengths and weaknesses and for planning instructional activities. Shepard’s chapter on classroom assessment,4 found in the same volume of Educational Measurement as that of Kane,1 provides an excellent discussion of this. And, just as you have outlined, it is generally sensible to combine information across modalities. This occasionally occurs for educational assessments – for example, in combining scores on disparate instruments to assign end-of-course or clerkship grades – though it may most commonly be done informally when a faculty member reviews test scores included in a portfolio of other information. I think it is unfortunate that most research has focused on the psychometric characteristics of individual tests or assessment methods. I like the emphasis that you and Cees [van der Vleuten] have placed on viewing individual assessments as part of an assessment system that should include a variety of (objective and subjective) components.5 It is relatively easy to determine the reproducibility of individual assessments; it is much harder to develop an assessment system that provides the broad range of information that would ideally be available, for both assessment of learning and assessment for learning. Yet this is hugely important because we are gaining more understanding of the way assessment influences students’ study and learning behaviour, both as a consequential effect and as a ‘washback’ effect.5–9 But in this we should not negate the direct effect that merely sitting a test (and having to remember the learned subject matter) has on retention of knowledge.10,11 Though pretending I have any understanding of medicine generally gets me into trouble, I think there are a few additional ‘tests’ in which somewhat unlike entities are combined: Apgar scores, total cholesterol and white blood cell counts with differential come to mind. These provide interesting examples to consider: clinical decisions can be informed by both the combined values and the component values, much as in educational tests. In passing, you alluded to the superiority of assessor judgement over checklists and I cannot resist commenting on that. I would probably agree with you if the purpose is formative: structuring assessments like the mini-clinical examination (mini-CEX) and case-based discussion so that trainees receive feedback from skilled assessors is likely to result in improved skills. For summative purposes, though, I think this is less clear, particularly for high-stakes, large-scale assessments. It is nearly impossible to place scores from different sites on comparable scales if judgements are highly subjective because the measurement error associated with assessors’ judgements (and patient mix if this is not controlled) is both large and confounded by the site where the assessment takes place. Checklists tend to be more transportable, but they need to be done well and not to trivialise the skills that are to be assessed. Well, it certainly does not get you in trouble here. I think your examples are great. I particularly like the Apgar score, because it is an example of qualitatively different aspects that are converted to numerical scores and then simply added. Yet the score is a good diagnostic test. For me, the bottom-line outcome of this correspondence is that we have to carefully consider how any combination of information is most meaningful. This may be a sort of qualitative expert(s) judgement, an algorithmic procedure or simple arithmetic, but – and here is where Kane1 comes in again – it has to be done in such a way that it produces the most plausible and defensible (or least falsified) argument. Simply applying a method because all the others are doing it will not suffice. In this respect, I came across a paper that I thought was a gem because it really made me think about the considerations that play a role when we are making the inference from observation to ‘score’.12 I have put score between apostrophes here because ‘score’ is any descriptor that is used to reduce the observable data into manageable packages. This may all seem a bit picky, but to me it is quite the contrary. From a programmatic view, for me there is no tension between assessment of and assessment for learning: they are both parts of the whole.13 Just as health care is based on screening and individual patient care, and the latter incorporates numerical, structured lab tests in conjunction with verbal unstructured judgements, the whole of assessment will have to include standardised testing, structured examinations and judgements. I am very happy that psychometricians (and I include you here) have conducted such good work in gaining a grasp on the quality of structured and standardised testing; now the onus is on others to do the same with the other components of assessment so that they can contribute credibly to the plausibility of their validity arguments. This seems like an excellent note to end on: the importance of taking a serious, balanced look at assessment that includes assessment for learning as well as assessment of learning, and thinking broadly about the validation of inferences drawn from assessment programmes as well as from individual assessments. Acknowledgements: none. Conflicts of interest: none. Ethical approval: not applicable. Contributors: this manuscript is a transcription of an original e-mail correspondence that took place between LS and DS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.010 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".