The Relationship between the Nature of Practice and Performance on a Cognitive Examination
Bibliographic record
Abstract
Certification by a specialty board affiliated with the American Board of Medical Specialties is one of the most widely used markers of physician competence in the United States.1 To enhance the meaning of this credential and in recognition of the need for periodic reassessments of physicians, the specialty boards have time-limited their certificates. To maintain certification, most of the boards have developed programs that incorporate (1) a check of credentials, (2) self-evaluation and/or continuing medical education, and (3) a secure written or computer-based examination. The secure examination is considered an integral component of certification because it provides assurance that physicians are keeping up with changes in medical knowledge and that they possess the ability to successfully manage patients' problems that are important but rarely encountered in practice. Moreover, from patients' and payers' perspectives, a secure examination lends more credibility to the certification process. Despite these benefits, the secure examination is contentious because many believe that it tests only the ability to recall factual knowledge and, as such, bears little relationship to the day-to-day practice of medicine.2 However, the growth of electronic records and databases has made it possible to begin to address this concern by comparing certification status and test performance with aspects of practice such as volume, process of care, and patients' outcomes. There is considerable evidence that physicians who treat large numbers of patients with a particular condition generally provide better care for such patients. Volume is directly associated with patients' outcomes, regardless of the discipline or procedure.3 Therefore, the relationship of practice volume to examination scores is an important part of any test-validation effort. A study involving the first geriatric medicine certifying examination indicated that the number of geriatric patients seen in practice was positively correlated with examination scores.4 Likewise, a study involving cardiologists indicated that performance on cardiovascular graphics questions was positively correlated with experience.5 Specifically, scores on the interpretation of echocardiograms were correlated with the numbers of echocardiograms interpreted in practice or training. Similarly, scores on the interpretation of arteriograms and ventriculograms were correlated with the numbers of angioplasties performed. More recently, a study of a cognitive recertification examination in critical care medicine showed that scores were related to the amounts of time physicians spent in the direct care of critically ill patients. This relationship persisted even after statistically removing performance on the initial certifying examination in the same discipline.6 There is also evidence that certification status and examination performance are related to the process of care. A study of physicians in Quebec compared consultation rates, inappropriate prescribing for the elderly, and mammography screening rates with licensing examination scores based partly on cognitive tests.7 Physicians with higher scores referred more patients for consultation, prescribed fewer inappropriate drugs and more disease-specific medications for symptom relief, and appropriately referred more women for mammography. Similarly, a study of certified and non-certified internists found differences in preventive care services favoring the certified physicians.8 Although it is difficult to measure good practice outcomes well, some progress is being made in using available data in the validation of cognitive examinations. A recent study investigated whether there were differences among certified and self-designated cardiologists, internists, and family practitioners in the mortality of their patients with acute myocardial infarction.9 Data for all myocardial infarctions for calendar year 1993 in Pennsylvania were analyzed. Certification was associated with a 15% reduction in mortality irrespective of specialty, and after taking account of severity of illness, hospital characteristics, patient volume, and years since graduation. Similarly, Ramsey and colleagues found differences in some outcomes that favored the certified physicians,8 and there are a few studies with similar positive results in other specialties.10,11,12 Given the significance of the topic and the need for additional investigation, the purpose of this study was to extend previous work by exploring in more detail the relationship between examination performance and the nature of practice. Specifically, candidates for recertification in critical care medicine supplied information, via a practice survey, about the amounts of time that they spent in the care of patients with cardiovascular and pulmonary problems (i.e., practice volumes). Moreover, they rated the complexity of the problems they saw. These practice data were compared with performances on the items from the examination that dealt specifically with cardiovascular and pulmonary problems to determine whether patient complexity, in addition to patient volume, was associated with test scores. Method Participants. The data are based on the candidates who attempted the 1997 and 1999 recertification examinations in critical care medicine and responded without error to the practice survey. All of these candidates had time-limited initial critical care certificates. Ninety-nine percent and 93% of their certificates expired in 1997 and 1999, respectively. In 1997, the average examinee had been certified in internal medicine in 1979 (18 years, SD = 4 years), and these physicians spent most of their time in direct patient care (mean = 70%, SD = 26%). In 1999, the examinee group's average candidate had been certified in internal medicine in 1982 (17 years, SD = 4 years), and these physicians also spent the majority of their time in direct patient care (mean = 72%, SD = 26%). Examinations. The 1997 and 1999 critical care medicine recertifying examinations each consisted of 120 single-best-answer questions, all of which were asked in the context of a clinical problem and most of which required synthesis and judgment to reach the correct response. Consistent with the purpose of the tests, the content focused on well-established principles of patient care that should be known without consulting medical resources. The questions were written by a test committee of experts, but before these questions were selected for the examinations, they were sent to critical care practitioners who rated them for relevance to practice. The examinations had average relevance ratings of more than 4 on a five-point rating scale, where 5 denoted “very relevant.” The same items appeared on the 1997 and 1999 critical care medicine initial certifying examinations. This study concentrated on the 1997 and 1999 exams' cardiovascular and pulmonary disease questions because problems in these areas were frequently encountered in practice by candidates, and they were the largest subsets of items on the examination. In addition, these examination years had substantial numbers of candidates taking the examination. Table 1 presents the numbers of items and their means (SD) for the 1997 and 1999 critical care medicine recertification examinations. For this study, subtest scores were reported on the raw score scale.TABLE 1: Descriptive Data and Regression Result for Physicians Who Took the 1997 and 1999 Critical Care Medicine Recertifying Examinations and Completed a Survey on Their Practices' CharacteristicsSince scores on initial certifying examinations are related to features of residency and fellowship training as well as fund of medical knowledge, the initial certifying examination in critical care medicine was used as a statistical control in this study.13 The analyses took account of the scores on this examination and attributed effects to other variables if they made independent contributions to the explanation of the performance. The scores on the initial certifying examination had been standardized against a national group with a mean of 500 (SD = 100) and were equated over years using a common-item linear equating technique. Table 1 presents the means and standard deviations for each cohort. Survey. When physicians applied for the examination, they were asked to supply information about their practices. Specifically, for each of the specialties of medicine, they were asked what percentage of time they spent with patients. Physicians whose responses did not amount to 100% were removed from the analysis. Candidates were also asked to rate on a five-point scale (where 1 was “not very complex” and 5 was “very complex”) the complexity of the cardiovascular and pulmonary disease cases they managed. Finally, to test the joint effect of time and complexity, the ratings were multiplied by percentage of time spent in an area. Table 1 presents descriptive data for these variables. Procedure. The data were submitted to four separate stepwise linear regressions, two (1997 and 1999) for cardiovascular disease and two (1997 and 1999) for pulmonary disease. The dependent measures were the cardiovascular and pulmonary subscores on the critical care medicine recertifying examination and the independent variables were (1) score on the initial critical care medicine certifying examination, (2) the frequency of patients' problems encountered in the area, (3) the complexity of those problems, and (4) the interaction of the two factors (i.e., frequency times complexity). Results In predicting the cardiovascular disorder subscore on the 1997 recertification examination, the critical care medicine certifying examination entered first into the regression equation (R2 change =.18, t = 10.44 p <.001), followed by the interaction of the frequency and complexity of patients with cardiovascular disorders (R2 change =.01, t = 2.72, p =.02). The other variables did not contribute significantly. For the 1999 recertification examination, similar results were obtained. The critical care medicine certifying examination again entered first (R2 change =.14, t = 7.74, p <.001), followed once more by the interaction of the frequency and complexity of patients with cardiovascular disorders (R2 change =.04, t = 3.98, p <.001). Again, the other variables did not contribute significantly. From these results we can infer that if there were critical care physicians who spent all of their time (100%) treating patients with complex cardiovascular problems they could be expected to perform 3.5 (1997) to 4 (1999) points better on cardiovascular disease items than would those who did not see any cardiovascular problems. There were not many physicians in this sample who spent all of their time treating such patients, but this constitutes a difference of 1.1 to 1.7 standard deviations. In predicting the pulmonary disorder subscore on the 1997 recertification examination, the critical care medicine certifying examination entered first (R2 change =.12, t = 8.20, p <.001), followed by frequency of patients with pulmonary disorders (R2 change =.02, t = 2.57, p <.001). The other variables did not contribute significantly. For the 1999 recertification examination, the critical care medicine certifying examination also entered first (R2 change =.23, t = 10.20, p <.001), followed by the complexity of patient problems (R2 change =.08, t = 5.16, p <.001), and frequency of patients with pulmonary disorders (R2 change =.01, t = 2.58, p =.01). Again, the other variables did not contribute significantly. From these results we can infer that if there were critical care physicians who spent all of their time treating patients with complex pulmonary problems they could be expected to perform.7 (1997) to 5.2 (1999) points better on pulmonary disease items than would those who did not see any pulmonary problems. There were not many physicians in this sample who spent all of their time treating such patients, but this constitutes a difference of between.4 to 1.6 standard deviations. Discussion The purpose of this study was to extend previous work by exploring the relationship between test performance and the nature of practice. Physicians who were recertifying in critical care medicine in 1997 and 1999 supplied information about the amounts of time they had spent in the care of patients with cardiovascular and pulmonary problems and the complexity of the problems they saw. These practice data were compared with performances on the relevant items from the examination. For cardiovascular diseases, the interaction between volume and complexity had a significant relationship with test scores for both years of the study, even after controlling for previous examination performance. For pulmonary diseases, only volume was a significant predictor in 1997, but both volume and complexity were significant in 1999. The magnitude of the effects was noteworthy, ranging from.4 to 1.7 standard deviations. These results should be interpreted with care because this study has several limitations. First, the number of questions in each content area was relatively small, and this attenuated the correlations that were reported. Second, the estimates of time and complexity were based on self-reported data, and a number of physicians made errors in filling out the form. Patients' records would clearly be a more accurate and less biased source of these data. Third, critical care medicine is a relatively new discipline and the vast majority of diplomates are certified in pulmonary disease. Therefore, these results may not generalize to a more homogeneous, less cross-disciplinary field. Despite these limitations, the results of this study replicate previous work indicating that cognitive examination test scores are associated with patient volume, and by implication from other studies, with patient outcomes. The study also found that there is a relationship between scores and the complexity of problems physicians see in practice. This finding bears more investigation, but it seems sensible that a practice that includes the challenge of treating many complex patients should lead to more knowledge and better judgment on the part of the physician. These findings, taken together with previous work, suggest that performance on a cognitive examination is related to performance in practice. Of course, this type of examination is not a substitute for rigorous evaluation of practice outcomes, nor is it broad enough to include important aspects of competence such as communication skills and professionalism. Nevertheless, until better measures are available for high-stakes use, the cognitive examination is a reasonable alternative.14 When such measures become available, there will still be a place for cognitive assessment of new developments in medicine and for patients' problems that are important but infrequently encountered in practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".