On the case for an interview in medical student selection
Bibliographic record
Abstract
Every other biennial Ottawa Conference on Clinical Competence is held outside North America, and in 2008, the venue was Melbourne (Ozzawa 2008). The conference was arranged into plenary sessions and concurrent streams, one of which was related to the selection of medical students. During one of these sessions, a presenter from the Middle East was questioned by a delegate from Europe about the importance of the medical student selection interview findings the former had presented. Notwithstanding the problems in language, there was also a schism in understanding. The questioner sought data on the specificity and sensitivity of the interview and its positive and negative predictive power. In our opinion, an undue focus on these aspects is both naive and misses the point. The utility of an interview as part of medical student selection is uncertain but certainly worthy of debate. Our medical school has used an interview, and at times a multifaceted interview process, as part of the selection mechanism since the first students were recruited at the end of 1967. All three of us have been subject to this system and have strong, both good and bad, memories of the event. As such, we have a bias that needs to be declared a priori. The first dean of the school reportedly introduced an interview at the outset to identify ‘bad buggers’. The modern view is that the interview should serve other wider purposes, although finding bad buggers is still desirable. The nature of the questioning from the European delegate at Ozzawa illustrates a commonly held perspective. Quantitative rationale brings comfort, and comfort is a sought-after commodity for medical student selection panels as the stakes are high and the competition is intense. The selection is rich in emotion for everyone involved, and the outcome is literally life-changing. The reality is that medical student selection defies simple quantitative expressions. A prediction has to be predictive of something. It would be nice if we could agree what is a good doctor and then see how well our selection processes predict the extent to which our graduates satisfy the agreed elements of ‘goodness’. As far as we can see, there are no tools yet to provide a quantitative estimate of a ‘good doctor’, and the very nature of an ethical doctor is intrinsically paradoxical; that is, we believe that a good doctor is a person who does the right thing on the one hand and is also someone who challenges systems, people and dogma and so on, on the other hand. The 20-year UK predictive study,1 which showed that the best predictor of ‘success’ for a doctor in their career was academic performance at school, used bald measures of success, such as ranking on medical school examinations, performance in house officer posts, time to achieve membership qualifications as well as publications and postgraduate qualifications; ‘A’ level results correlated positively with the first three outcomes. Given the data available to us about both previous academic performance and entering careers of choice and the misalignment of career choice by medical graduates and community need,2-4 it could be argued that academic performance at school might be a negative predictor of a graduate becoming a useful doctor. It is just as difficult to define a bad doctor in a way that facilitates statistical analysis; bearing in mind, we started this process 40 years ago to identify bad buggers. The likelihood, nature and outcome of a complaint during a medical career are highly dependent on factors that are outside a doctor’s control, such as the nature of the discipline and the medicolegal milieu in which he or she practises. Although a report of unprofessional behaviour as a medical student is a statistically significant risk factor for subsequent regulatory problems,5, 6 the positive predictive value in this context is very low. Based on our experience, we would be the first to admit that a structured interview alone does not detect all those with undesirable personality traits. However, it is the first element in a necessary process of successive observations by experienced faculty, over time and in a range of contexts. There are two other major problems with a quantitative approach to validating or refuting the utility of interviews in the selection of medical students. The first is that interviews almost always only occur for a subcohort of applicants who are identified by academic performance that exceeds some threshold. The number of applicants for medical schools, and the logistics, cost and time taken in interviews are such that some previous short-listing is essential. The end result is that the cohort coming to interview is increasingly ‘compressed’ in terms of both academic grades and the personal qualities that are commonly associated with doing well at school and university. A related consideration is the desire to have a diverse student cohort accepting that this has the best chance of producing doctors to work in areas of health need.6 One diversification strategy we use is to have separate admission processes for candidates of Maori and Pacific ethnicity (up to 30 places in a total domestic cohort of 155 students) and for rural origin students (20 places). The second is that interviews will always be subject to the bias of the interviewers and that these same biases will be influential in choosing the outcome measures; the evaluation then is self-fulfilling. The profession of medicine has long been appropriately criticised for just such guild-protective behaviour.7 We are fortunate to have an experienced and well-calibrated, yet diverse, interviewer pool from which to draw. In contrast, there is a strong qualitative argument for using an interview to help select medical students. This is especially true when the interview is also assessed qualitatively. With the exception of the Maori and Pacific students who have multiple mini-interviews, each candidate above the grade point average (GPA) threshold has a single 25-min structured interview over five domains with two interviewers. Each interviewer grades independently in the first instance; a final grade is then agreed. We struggled for a numeric concordance of greater than 70% between interviewers of the same applicant when we used visual analogue scales of domains such as communication, but this improved to almost 100% when we changed to ‘word-picture’ categorical outcomes. There are five main reasons why we would argue for the retention of an interview in medical student selection. The first is still to identify at least some of the bad buggers, and, also, hopefully, a few ‘good buggers’. At Auckland, those who score full marks from each interviewer are automatically accepted, and those scoring minimum marks, rejected, regardless of their academic rankings; accepting that they have already reached the threshold for interview. In effect, this changes few of the places offered. The second is that the interview provides some applicants an opportunity to ‘self-destruct’ without loss of face and in a way that avoids conflict with parents, significant others, teachers and so on, many of whom might have a clear, but, unilateral ambition for the person to become a doctor. Every year, we see a number of such self-destructs. The interview is often the first and only opportunity for a naturally gifted young person to opt out of medicine; for example, since a day early in their secondary school life when a teacher remarked to their parents that they were ‘bright enough to do medicine’. Third, an interview can contribute to constructive social engineering. In our opinion, medical schools have a social contract with the community in which they are based and by which they are funded to train an appropriate medical workforce. This requires a process that, at least, minimizes maldistributions. We have classified such maldistributions as being variously disciplinary, cultural and demographic.8, 9 The ethnic demography of the Auckland region is very dissimilar to the student cohort, which would have arisen over the last 5 years if selection had been based on GPA alone. In addition to GPA (60% weighting), student selection at our medical school is also determined by interview categorization (25% weighting – if not an automatic ‘accept’ or ‘reject’ as outlined earlier) and aptitude testing (Undergraduate Medicine and Health Sciences Admission Test (UMAT), © Australian Council for Educational Research) (15% weighting).10 It is not our intent to critique the UMAT here or suggest what the scores actually mean, but less than 5% of the variance of our selection process is determined by these scores. That is, for our applicants, UMAT seems to be influenced by the same parameters as the GPA. A review of the actual ethnic mix of our medical students shows that the interview has an unintended, but arguably positive, outcome of a much closer alignment of the ethnicities of our student population to that of our community. Fourth, and again considering the social contract described above, the interview has a high level of face validity with both health consumer and provider communities. Australians and New Zealanders are strongly inclined to meritocracies such that the weighted balloting used by some European medical schools is unlikely to have appeal.11 Evaluation interviews with consumer community stakeholders have shown a strong preference for a medical student selection process that is not based on academic performance alone. Similarly, doctors are also very interested in being part of a process that determines the future of the profession. This concept of face validity is not susceptible to quantitative summary. Fifth, and finally, the interview has an implicit formative nature and may be seen as the first step in the professional development of the doctor of the future. At Auckland, we are committed to maintaining an interview in our medical student selection process and for what we see as good reasons. We do not expect to be able to convince sceptics that this is a good idea by any employment of statistics, and we recognize that by the time tracking systems identify problems, generational and other differences will make any review of extant selection practice somewhat after the event and irrelevant. This is not an argument against such tracking but rather is a call for a pragmatic and holistic view of how we go about shaping the future medical workforce.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.073 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.006 |
| Insufficient payload (model declined to judge) | 0.022 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".