MétaCan
Menu
Back to cohort

Certification and Recertification

2000· article· en· W2079113204 on OpenAlexaffabout
John Cunnington, Geoff Norman

Bibliographic record

VenueAcademic Medicine · 2000
Typearticle
Languageen
FieldMedicine
TopicInnovations in Medical Education
Canadian institutionsMcMaster University
Fundersnot available
KeywordsCertificationMedical educationMedicinePolitical scienceLaw

Abstract

fetched live from OpenAlex

Are certification and recertification the same? If not, what is the difference? Should there be one standard for both certification and recertification, and if so, should there be only one test for both? These are the conceptual questions that must be addressed by any organization responsible for certification or licensure. Any attempt to answer these questions must start from the position that a professional group must subscribe to a basic minimum standard that all practitioners—novice and seasoned—must meet. Standards at this level are not negotiable. (It is also worth saying that because the standards are not defined in words and spelled out on paper does not mean that the profession does not have standards; medicine, like other professions, has unspoken standards that become most apparent when they are violated.) But the question of whether the same test should be used for both certification and recertification is a different one. The purpose of certification is to assure society that a physician entering practice has demonstrated the profession's basic minimum requirements for knowledge, skills, behaviors, and judgment, which, for the sake of brevity, we call competency. When faced with certifying or licensing applicants they have not trained themselves, organizations may use reports generated at the training institution, but typically (in North America at least) they rely upon their own examinations, which serve to provide an external guarantee (as well as a seal of approval) of the candidates' competency. Fundamentally, such a test can consist of nothing more than a series of sample questions or “biopsies” from the field of interest. Certainly, the test design must be broad enough to adequately survey the requisite knowledge, skills, and behaviors, and it is the task of the test designers to decide how many test questions are required for the certifying organization to be confident in certifying competence. As only so much information can be gathered in a unit of time, such tests typically require several days to complete and are usually composed of at least one or two written papers, and, increasingly, some form of performance-based assessment, such as actual or simulated patient encounters. While applicants cannot be tested on all possible clinical problems, the expectation is that the sample of problems on the test will be broad enough for the certifying organization to have a high degree of confidence that all important domains and basic standards have been covered. It is at this point that recertification differs. Recertification starts from a position of confidence that the candidate has already met the high and rigorous standards of the certifying body. The question, then, is not is this candidate competent to enter practice, but rather has the candidate kept up his or her previously demonstrated skills to a sufficient level that the certifying body can justifiably assure public confidence in his or her performance? We argue that testing to achieve this goal need not be as extensive, on three grounds: (1) content validity, (2) reliability, and (3) the relative cost of wrong decisions. HOW RECERTIFICATION TESTING DIFFERS FROM CERTIFICATION TESTING Content Validity Because a recertification test is designed to determine whether a physician has maintained previously demonstrated skills, it need not cover all content domains as thoroughly as would a certification test. Reliability Simply put, it is the job of any test to discriminate between those candidates who have high levels of performance and those who have low levels of performance. A certifying examination is usually taken by many relatively similar candidates, who by virtue of recent and similar training are relatively similar in their knowledge, skills, and behaviors. To reliably discriminate among them requires relatively large number of questions. By contrast, recertification typically involves looking at the skills of those who have been in practice for ten, 20, or even 30 years. The passage of time, plus differences in local facilities and practice patterns, results in a highly diverse group.1 A test with fewer questions is sufficient to assess performance in this more heterogeneous group. Decision Making Any instrument to assess competence can be regarded as a diagnostic test to detect a relatively rare disease called “incompetence.” However, in diagnosing incompetence we never have available a pathologist who represents the “gold standard.” Nevertheless, we can still conceptualize two distributions of test scores, one corresponding to truly incompetent physicians, and a second, higher, distribution corresponding to truly competent physicians. As with any test, there will be true negatives (competent physicians called competent), true positives (incompetent physicians called incompetent), false negatives (incompetent physicians called competent), and false positives (competent physicians called incompetent). Now, for every test, we must establish a cutoff score, below which we will declare an individual incompetent and above which we will declare him or her competent. Depending on the location of the cutoff score, we will run a greater risk of false-positive or false-negative results, but there will always be miscalls. While we may wish to act as if the test were perfect, it cannot be; our intention should be to minimize errors, but they cannot be altogether eliminated. Of course, there is a cost associated with each miscall, and this varies by situation. In licensure decisions, the cost of a false negative amounts to giving an incompetent individual a license to practice medicine for the next 40 years; the cost of a false positive is that the individual affected must reapply and probably will pass the next year. So for a national licensing body, such as the National Board of Medical Examiners in the United States or the Medical Council of Canada (MCC), it makes sense to set the pass mark relatively higher and reduce the likelihood of a false negative. Conversely, for recertification, the cost of a false-positive is the possibility that a physician who has performed well for many years and is still functioning well may be coerced into giving up his or her practice, which will have major adverse effects upon the physician's life and patients. Of course the cost of a false negative is that an incompetent physician who may pose a risk to patients remains in practice. However, considering the relative costs under these circumstances, it is reasonable to set the hurdle lower—the experienced physician must be given the benefit of the doubt. THE EVIDENCE So far this discussion has been hypothetical. We certainly have some evidence that the recertification cohort is more heterogeneous and hence easier to distinguish among than are the new certificants. Norman et al.1 showed that a relatively brief recertification test could achieve adequate reliability in discriminating among such physicians. In particular, a two-hour performance examination using three to five standardized patients had a reliability of about 0.9. Conversely the 20-station licensing OSCE of the MCC, which requires a full day of testing, has a reliability of only 0.6–0.78.2 We have no direct evidence that either certification or recertification groups deliberately adjust the cutoff points higher or lower to accommodate different relative costs. But there is at least some circumstantial evidence that such an action may be operative. For the MCC, the population is drawn, by and large, from recent graduates, who tend, by most measures, to be at the peak of their competence. Nevertheless, the MCC fails 4.5% of them. In contrast, data from performance-assessment programs of at least one organization that assesses physicians in practice gives another picture. The College of Physicians and Surgeons of Ontario (CPSO) has a two-part approach to assessing practicing. The peer-assessment program does random peer-assessed, practice-based audits.3 About 10% of primary care physicians either are found to be functioning at unacceptably low levels or their records are so deficient that the assessors are unable to determine at what levels they are functioning. As primary care physicians compose about half the practicing physicians in Canada and have a much higher rate of identified deficiency, we could say that this group represents the bottom 5% of all practicing physicians. Those physicians whose performances are identified by peer review as having possible deficiencies are referred for further assessment to the Physician Review Program (PREP). PREP is a one-day standardized assessment that uses multiple-choice questions, standardized-patient review, and chart-stimulated recall to evaluate the knowledge, skills, and behaviors of primary care physicians.4 Of the physicians referred to PREP, only about 10%—that is, about 0.5% of all physicians—are rated as unsafe to practice. Thus, leaving aside the issue that older physicians, who make up a significant proportion of the identified physicians, must on average be practicing at lower levels than recent graduates (and should therefore have a higher failure rate), the CPSO has a failure rate that in some absolute sense is about one tenth that of the MCC. The roughly tenfold lower failure rate for physicians taking the recertification examination, as opposed to the failure rate of those taking the certification examination, suggests that the CPSO is aware of the risk of falsely identifying a physician as incompetent and has deliberately set its cutoff point high to minimize such error. CONCLUSION Achieving a defensible approach to recertification is a long and difficult process. On the one hand, such efforts must protect society from incompetent and potentially dangerous practitioners, but on the other hand they must provide adequate safeguards to prevent practitioners' having their careers and livelihoods inappropriately disrupted. A clear understanding of the values inherent in the decision-making process may lead to more effective, efficient, and defensible approaches to recertification.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.016
metaresearch head score (Gemma)0.085
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.984
Threshold uncertainty score0.169

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0160.085
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.002
Science and technology studies0.0020.007
Scholarly communication0.0070.007
Open science0.0020.005
Research integrity0.0040.005
Insufficient payload (model declined to judge)0.0500.018

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.051
GPT teacher head0.372
Teacher spread0.321 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainEvaluation
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2000
Admission routes2
Has abstractyes

Explore more

Same venueAcademic MedicineSame topicInnovations in Medical EducationFrench-language works237,207