Modified endocrinology script concordance test: evaluating the reliability and construct validity for assessing clinical reasoning
Bibliographic record
Abstract
INTRODUCTION A 35-year-old woman with a history of type 1 diabetes mellitus was admitted for symptoms of fever and vomiting which started 3 days ago. The attending endocrinology resident was concerned about the patient’s persistently low blood pressure despite resuscitation with large volumes of intravenous isotonic saline. Her consultant supervisor identified areas of brown pigmentation on the patient’s palm during physical examination and noted a low plasma sodium of 128 mmol/L. He promptly asked the team to draw blood for random cortisol measurements before starting intravenous hydrocortisone for a presumed diagnosis of hypocortisolism. The patient’s blood pressure promptly improved within the next hour, while returned results confirmed the diagnosis. As described in the case vignette, experienced clinicians match clinical findings against a mental bank of organised schemas of clinical symptoms and expected signs known as illness scripts.[1] The degree of matching between patients’ clinical findings and the scripts guides the doctor in deciding an appropriate set of investigations required to arrive at the most likely diagnosis and treatment.[2] Senior doctors develop a rich database of well-organised illness scripts after gaining experience in years of clinical practice. Every clinical clue leads to the activation of a predefined illness script, which prompts the physician to find evidence to support or refute the hypothesis. Conversely, junior doctors, who are in the formative years of medical training, require more time to organise these pieces of information within the clinical puzzle. Are there assessment methods to evaluate how the quality and use of illness scripts among junior doctors change with training and how relevant such methods are in the field of endocrinology? The script concordance test (SCT) is a validated assessment tool that judges clinical reasoning skills in uncertain clinical situations.[3] Performance in the assessment is benchmarked against a panel of clinical experts. Candidates have to interpret data within written case scenarios under circumstances of uncertainty, before deciding the likelihood of a diagnosis or appropriateness of an investigation or treatment based on the information added to a clinical stem.[4] Current evidence supports SCT as a more valid measure of clinical reasoning in the Internal Medicine field in comparison to the traditional multiple-choice questions, where candidates are asked to choose a single best answer (SBA) based on factual recall.[5] While the practice of clinical endocrinology frequently occurs amidst uncertain conditions described within an SCT, no specific assessments have been created to assess clinical reasoning in trainees practising within this field. The SCT has been applied in the assessment of clinical reasoning across various medical specialities in both undergraduate and postgraduate set-ups.[6–10] Within the specific field of endocrinology, attempts to utilise SCT have only been described among medical undergraduates within a problem-based learning curriculum[11] and in the training of pharmacy students in the areas of diabetes pharmacology.[12] This study sought to evaluate the reliability and construct validity of an endocrinology SCT in assessing the clinical reasoning skills among senior resident trainees across two training institutions in Singapore as compared to the performance of an expert panel of 15 consultant endocrinologists. Specifically, the study aimed to determine if clinical reasoning among endocrinology speciality trainees, as assessed by performance in SCT, correlates with the amount of time spent in postgraduate residency training. In addition, the SCT was also modified to determine if junior trainees make random selection that is not backed by reasoning (‘wild guess’) more frequently than senior trainees as a sign to further justify the construct validity of the test. The study also attempted to evaluate participants’ perception of how SCT influenced their learning of clinical reasoning. METHODS This was a prospective, multicentre study involving 17 endocrinology speciality trainees (senior residents), resident physicians and recent speciality training graduates (first-year associate consultants) working in two training institutions in Singapore. The senior residents are required to undergo a 3-year training programme (in addition to 3 years of foundation training in Internal Medicine) before they are assessed and formally accredited as an endocrinologist. The two institutions involved in the study are registered training sites within the residency system of postgraduate medical training in Singapore. Training is conducted based on a common curriculum managed by the Singapore Residency Advisory Committee and the Accreditation Council for Graduate Medical Education-International (ACGME-I). The first author (Mok SF) created a set of endocrinology SCT comprising 15 clinical stems with 64 distinct questions. The questions were based on common endocrinology conditions (e.g. diabetes mellitus, thyroid dysfunction, pituitary and adrenal disorders, electrolyte abnormalities) that trainees need to be familiar with. The construction of SCT was guided by recommendations from subject matter experts.[4] The SCT questions were vetted by a co-author (Seow CJ) before they were administered to the expert panel and trainees. Brief explanation of construct and principles of SCT Figure 1 shows an example of an SCT created based on the subject of evaluating a patient with clinical features of thyrotoxicosis and deranged thyroid function test. Each question begins with a case vignette that provides adequate clinical context, but still leaves a fair amount of uncertainty, so as to mimic realistic clinical scenarios. Participants then respond to a series of questions by rating the likelihood or appropriateness of a diagnosis, investigation or treatment option on a five-point Likert scale.Figure 1: Sample question demonstrating the clinical stem and the follow-up questions seeking judgements about the likelihood of a diagnosis when added information is found on history or physical examination of the same patient. Responders are asked to rate the likelihood of the option (aetiology of thyrotoxicosis in this example) using an annotated 5-point Likert scale. Responders are also asked to indicate if they are making a random guess for the particular question by marking a chose in the final column with the red question mark. In this instance, the responder chose a score of ‘−1’ for question (b), indicating that the physical finding of cervical lymphadenopathy made the diagnosis of toxic thyroid adenoma less likely. This was, however, a random guess and not based on actual clinical reasoning, and hence he went on to mark a cross in the right-most column with the question mark heading.The SCT created was first administered to 15 selected consultants in the two training institutions in the pre-study phase at end of 2019. These consultants are endocrinologists who are certified by the local Specialist Accreditation Board, and their duration of practice as a specialist ranged from 3 to 22 years. They formed the expert panel whose performance was analysed according to the recommended methods to generate reference scores.[4] Briefly, the modal response (i.e. option selected with the highest frequency among consultant respondents) was allocated 1 point, while the other options were accorded proportionately fewer marks. For example, if ten experts performed the test and eight individuals selected the ‘+1’ option while two individuals selected the ‘+2’ option, the score attributed to ‘+1’ would be 8/8 or 1, and the score attributed to the ‘+2’ option would be 2/8 or 0.25. The ‘−2’, ‘−1’ and ‘0’ options for that specific question would all yield a score of 0. All consultants only attempted this SCT once. The described method was employed for all questions in the SCT to determine the scoring weightage for each option. Calculation of SCT score was performed with a spreadsheet-based calculator made available by the University of Montreal’s School of Medicine, Canada (cpass.umontreal.ca). This calculator (embedded with the weightage based on the expert panel’s response) was then used to determine the scores of individual study participants. Seventeen endocrinology junior doctors were invited to participate in the study after the local Domain Specific Review Board (DSRB) granted permission in January 2020 for the study to commence (DSRB reference number 2019/00882). Both consultants in the expert panel and trainee participants were allowed 60 min to complete the test. Participants’ responses in the SCT were anonymised, and only individual year of training was indicated. Assessment of utility of modified SCT in the assessment of clinical reasoning The reliability and degree of internal correlation of the SCT were assessed via the Cronbach’s alpha metric. The construct validity was determined by comparing the mean score of the expert panel to that of the trainees. In addition, participants’ performance in the SCT was also analysed to check for correlation with the year of training as an added evidence of construct validity. Specifically, correlation coefficients were generated to analyse the relationship between individual SCT score and the frequency of declared guesses with the variable of year of training. All quantitative data analyses were carried out with IBM SPSS Statistics version 20.0 for Windows (IBM Corp, Armonk, NY, USA). A P-value < 0.05 was used to denote statistical significance in this study. All participants were invited to attend a posttest debrief to understand the explanation underpinning the modal answers for each SCT question. To understand attendees’ perceptions on whether the conduct of the study influenced their learning of clinical reasoning, they were asked to provide written, anonymous feedback to the following two questions via an electronic survey: “How do you compare SCT to a standard multiple-choice question in its impact on your development of clinical reasoning?” and “Please kindly indicate how, if at all, exposure to this endocrinology SCT will change the way you reason and arrive at diagnoses or management plans”. The responses were thematically analysed by the authors to check for common themes. RESULTS The 17 junior doctors completed their anonymised SCTs between January 2020 and February 2020. The distribution of the study participants’ demographic data and scores are indicated in Table 1.Table 1: Study participants’ demographic data and scores.Cronbach’s alpha for the SCT, as determined by the SCT calculator, measured 0.62, suggesting fair reliability. The mean score from the expert panel (78.8 ± standard deviation [SD] 7.3) was significantly higher than that of the trainees (70.7 ± 5.5) by unpaired t-test (mean difference = 8.0, P = 0.0018). Trainees’ year of training showed moderate positive correlation with individual SCT score (R = 0.557, P = 0.0202) and strong negative correlation with the frequency of guessing (R = −0.740, P = 0.0007). These findings suggested strong construct validity of our novel SCT. A review of the frequency of guesses also showed that there was clustering around specific topics such as approach to elevated free thyroxine and thyrotrophin (thyroid stimulating hormone [TSH]) and localisation in Cushing syndrome. These were deemed to be more advanced and complex topics that were less familiar to novice trainees in their first and second year of training. There was, however, a wide degree of variation in the frequency of guessing (range 0–22, mean 5.7 ± 7.1) among the 17 participants. Eight of the 17 SCT responders attended the posttest debrief, which was conducted via digital conferencing in April 2020 due to the need for physical distancing during the coronavirus disease 2019 (COVID-19) pandemic. Analysis of the qualitative comments indicated that trainees favoured SCT over traditional SBA questions, as they could focus on determining how new pieces of information strengthen or weaken their diagnostic hypothesis in comparison to their own illness scripts. The SCT was also favoured because responders could appreciate experts’ reasoning behind the choice of the modal answer during the debrief session. DISCUSSION In this study, we created a novel and modified SCT in the speciality domain of endocrinology to assess clinical reasoning in postgraduate trainees. To the best of our knowledge, this is the first time an SCT has been applied in this clinical field. Through comparisons between experts’ and trainees’ mean score and correlation between individual scores and their training year and frequency of guessing, we demonstrated the construct validity of this assessment via a pilot study. Significantly, the responders expressed that exposure to this SCT has given them a new-found appreciation of how to process information when approaching uncertain clinical situations, so as to judge the likelihood and appropriateness of evaluation and treatment methods. Through practice and experience accumulated from years of training, senior doctors develop more mature scripts that can be recalled rapidly and become more adept at making diagnoses amidst ambiguity. This likely accounted for the expert panel’s significantly superior mean score compared to that of trainee responders in our study. Illness script development and maturation is also known to improve with time in training,[3] and this was exemplified by more senior trainees having higher SCT scores and lower frequency of guessing. Based on the current ACGME-I framework of competency-based approach to postgraduate medical training in Singapore, medical knowledge and patient care are the two domains most closely related to the area of clinical reasoning. Traditional knowledge-based assessments, such as the endocrinology self-assessment programme, are more suited for evaluating the ability of trainees to recall medical science-related facts, but may not determine their ability to apply such knowledge in actual clinical practice. Workplace-based assessments, such as the chart-stimulated recall and mini-Clinical Evaluation Exercise, are typically used within residents’ portfolio, but are limited by the variability of assessors and clinical scenarios, which reduce their reliability, while adoption is also challenged by time pressure within real-world practice.[13,14] To overcome the limitations of the existing assessment framework, SCTs can be deployed to provide more reliable standardised testing. The use of realistic clinical stems embedded with controlled degree of ambiguity helps to assess examinees’ performance in real-world medical practice in a surrogate manner while preserving assessment validity. This, however, will require more SCTs to be created for the purpose of repeat routine testing on a scheduled basis and may add to the time burden of assessment. Despite the positive findings from our study, there were several limitations that should be addressed. Our SCT’s reliability was only fair, given the relatively low Cronbach’s alpha score compared to a desired value of 0.75.[4] This may be due to the lower number of unique questions within our SCT (64), which fell short of the recommended minimum of 75,[15] after our initial vetting led to the removal of several test questions. In addition, our SCT consisted of several different endocrinology topics that may not correlate well; a trainee strong in applying clinical reasoning in a topic related to diabetes mellitus may, however, be weaker in the areas of pituitary disease. This is also observed in the clustering of guesses in more challenging clinical domains within our SCT. Future iterations may need to involve the creation of distinct SCTs for individual endocrinology subspeciality topics. Our case for construct validity of the modified SCT relied on the observed difference in mean test scores between only two groups of subjects: the expert panel and postgraduate trainees. Two local studies evaluating the use of SCTs in the field of neurology and pulmonary/critical care medicine applied their respective tests in both trainees and undergraduate students[8,16] in comparison to experts. This provided more robust data to appraise the construct validity of SCT in these contexts since experts performed better than postgraduate learners, who outperformed undergraduate students in those studies. Such modification may be adapted in a future extension of our current study to further evaluate our SCT’s validity. In addition, the response format of the SCT itself may threaten assessment validity and reduce the ability to distinguish between trainees with strong and weak clinical reasoning skills. Examinees can opt to ‘play it safe’ and select the middle option (0: neither more likely nor less likely) to collect some points since there is a high likelihood of expert choosing to sit on the fence when faced with an uncertain situation.[17] Previous studies have described employing techniques of making trainees indicate their reason for selecting their chosen options (‘Think Aloud’) as a method to circumvent the problem of guessing bias.[18,19] Our group utilised a novel method of having SCT trainee responders declare their tendency to guess (as a surrogate of degree of perceived uncertainty), so as to overcome this inherent limitation. While there appears to be negative correlation between participant’s year of training and guessing frequency, this is, however, dependent on how honest responders elect to be when attempting the SCT; hence, there is a definite risk of underreporting of guessing tendencies. We acknowledge that the use of declared guessing to control for guessing bias is not supported by the existing literature, while individual perception of uncertainty is also highly varied. These factors may hence reduce the construct validity of the modified SCT within the study. Finally, the study was unable to evaluate the impact of the SCT on the participants’ actual clinical reasoning performance. Only limited subjective qualitative comments were provided by a small number (eight out of 17) of SCT responders to indicate their perception of the impact of the SCT on their learning of clinical reasoning. The conduct of the survey after SCT debriefing would also have biased the participants’ responses, as they would have considered the debriefing as part of the SCT itself. A more in-depth and unbiased analysis of the perceived educational impact of the SCT could have been obtained via individual interviews or focused group discussions from both the consultants within the expert panel and the junior doctor participants immediately after the SCT. In addition, more robust methods (e.g. repeat SCT testing, use of other workplace-based assessments during clinical practice) will need to be employed to longitudinally track how, if at all, the SCT influenced actual clinical reasoning in our trainees. In conclusion, our novel SCT has demonstrated moderate construct validity and fair reliability in the evaluation of clinical reasoning among endocrinology postgraduate trainees within the confines of an assessment pilot. Financial support and sponsorship Nil. Conflicts of interest There are no conflicts of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.647 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".