MétaCan
Menu
Retour à la cohorte
Enregistrement W2079113204 · doi:10.1097/00001888-200006000-00011

Certification and Recertification

2000· article· en· W2079113204 sur OpenAlexaffabout
John Cunnington, Geoff Norman

Notice bibliographique

RevueAcademic Medicine · 2000
Typearticle
Langueen
DomaineMedicine
ThématiqueInnovations in Medical Education
Établissements canadiensMcMaster University
Organismes subventionnairesnon disponible
Mots-clésCertificationMedical educationMedicinePolitical scienceLaw

Résumé

récupéré en direct d'OpenAlex

Are certification and recertification the same? If not, what is the difference? Should there be one standard for both certification and recertification, and if so, should there be only one test for both? These are the conceptual questions that must be addressed by any organization responsible for certification or licensure. Any attempt to answer these questions must start from the position that a professional group must subscribe to a basic minimum standard that all practitioners—novice and seasoned—must meet. Standards at this level are not negotiable. (It is also worth saying that because the standards are not defined in words and spelled out on paper does not mean that the profession does not have standards; medicine, like other professions, has unspoken standards that become most apparent when they are violated.) But the question of whether the same test should be used for both certification and recertification is a different one. The purpose of certification is to assure society that a physician entering practice has demonstrated the profession's basic minimum requirements for knowledge, skills, behaviors, and judgment, which, for the sake of brevity, we call competency. When faced with certifying or licensing applicants they have not trained themselves, organizations may use reports generated at the training institution, but typically (in North America at least) they rely upon their own examinations, which serve to provide an external guarantee (as well as a seal of approval) of the candidates' competency. Fundamentally, such a test can consist of nothing more than a series of sample questions or “biopsies” from the field of interest. Certainly, the test design must be broad enough to adequately survey the requisite knowledge, skills, and behaviors, and it is the task of the test designers to decide how many test questions are required for the certifying organization to be confident in certifying competence. As only so much information can be gathered in a unit of time, such tests typically require several days to complete and are usually composed of at least one or two written papers, and, increasingly, some form of performance-based assessment, such as actual or simulated patient encounters. While applicants cannot be tested on all possible clinical problems, the expectation is that the sample of problems on the test will be broad enough for the certifying organization to have a high degree of confidence that all important domains and basic standards have been covered. It is at this point that recertification differs. Recertification starts from a position of confidence that the candidate has already met the high and rigorous standards of the certifying body. The question, then, is not is this candidate competent to enter practice, but rather has the candidate kept up his or her previously demonstrated skills to a sufficient level that the certifying body can justifiably assure public confidence in his or her performance? We argue that testing to achieve this goal need not be as extensive, on three grounds: (1) content validity, (2) reliability, and (3) the relative cost of wrong decisions. HOW RECERTIFICATION TESTING DIFFERS FROM CERTIFICATION TESTING Content Validity Because a recertification test is designed to determine whether a physician has maintained previously demonstrated skills, it need not cover all content domains as thoroughly as would a certification test. Reliability Simply put, it is the job of any test to discriminate between those candidates who have high levels of performance and those who have low levels of performance. A certifying examination is usually taken by many relatively similar candidates, who by virtue of recent and similar training are relatively similar in their knowledge, skills, and behaviors. To reliably discriminate among them requires relatively large number of questions. By contrast, recertification typically involves looking at the skills of those who have been in practice for ten, 20, or even 30 years. The passage of time, plus differences in local facilities and practice patterns, results in a highly diverse group.1 A test with fewer questions is sufficient to assess performance in this more heterogeneous group. Decision Making Any instrument to assess competence can be regarded as a diagnostic test to detect a relatively rare disease called “incompetence.” However, in diagnosing incompetence we never have available a pathologist who represents the “gold standard.” Nevertheless, we can still conceptualize two distributions of test scores, one corresponding to truly incompetent physicians, and a second, higher, distribution corresponding to truly competent physicians. As with any test, there will be true negatives (competent physicians called competent), true positives (incompetent physicians called incompetent), false negatives (incompetent physicians called competent), and false positives (competent physicians called incompetent). Now, for every test, we must establish a cutoff score, below which we will declare an individual incompetent and above which we will declare him or her competent. Depending on the location of the cutoff score, we will run a greater risk of false-positive or false-negative results, but there will always be miscalls. While we may wish to act as if the test were perfect, it cannot be; our intention should be to minimize errors, but they cannot be altogether eliminated. Of course, there is a cost associated with each miscall, and this varies by situation. In licensure decisions, the cost of a false negative amounts to giving an incompetent individual a license to practice medicine for the next 40 years; the cost of a false positive is that the individual affected must reapply and probably will pass the next year. So for a national licensing body, such as the National Board of Medical Examiners in the United States or the Medical Council of Canada (MCC), it makes sense to set the pass mark relatively higher and reduce the likelihood of a false negative. Conversely, for recertification, the cost of a false-positive is the possibility that a physician who has performed well for many years and is still functioning well may be coerced into giving up his or her practice, which will have major adverse effects upon the physician's life and patients. Of course the cost of a false negative is that an incompetent physician who may pose a risk to patients remains in practice. However, considering the relative costs under these circumstances, it is reasonable to set the hurdle lower—the experienced physician must be given the benefit of the doubt. THE EVIDENCE So far this discussion has been hypothetical. We certainly have some evidence that the recertification cohort is more heterogeneous and hence easier to distinguish among than are the new certificants. Norman et al.1 showed that a relatively brief recertification test could achieve adequate reliability in discriminating among such physicians. In particular, a two-hour performance examination using three to five standardized patients had a reliability of about 0.9. Conversely the 20-station licensing OSCE of the MCC, which requires a full day of testing, has a reliability of only 0.6–0.78.2 We have no direct evidence that either certification or recertification groups deliberately adjust the cutoff points higher or lower to accommodate different relative costs. But there is at least some circumstantial evidence that such an action may be operative. For the MCC, the population is drawn, by and large, from recent graduates, who tend, by most measures, to be at the peak of their competence. Nevertheless, the MCC fails 4.5% of them. In contrast, data from performance-assessment programs of at least one organization that assesses physicians in practice gives another picture. The College of Physicians and Surgeons of Ontario (CPSO) has a two-part approach to assessing practicing. The peer-assessment program does random peer-assessed, practice-based audits.3 About 10% of primary care physicians either are found to be functioning at unacceptably low levels or their records are so deficient that the assessors are unable to determine at what levels they are functioning. As primary care physicians compose about half the practicing physicians in Canada and have a much higher rate of identified deficiency, we could say that this group represents the bottom 5% of all practicing physicians. Those physicians whose performances are identified by peer review as having possible deficiencies are referred for further assessment to the Physician Review Program (PREP). PREP is a one-day standardized assessment that uses multiple-choice questions, standardized-patient review, and chart-stimulated recall to evaluate the knowledge, skills, and behaviors of primary care physicians.4 Of the physicians referred to PREP, only about 10%—that is, about 0.5% of all physicians—are rated as unsafe to practice. Thus, leaving aside the issue that older physicians, who make up a significant proportion of the identified physicians, must on average be practicing at lower levels than recent graduates (and should therefore have a higher failure rate), the CPSO has a failure rate that in some absolute sense is about one tenth that of the MCC. The roughly tenfold lower failure rate for physicians taking the recertification examination, as opposed to the failure rate of those taking the certification examination, suggests that the CPSO is aware of the risk of falsely identifying a physician as incompetent and has deliberately set its cutoff point high to minimize such error. CONCLUSION Achieving a defensible approach to recertification is a long and difficult process. On the one hand, such efforts must protect society from incompetent and potentially dangerous practitioners, but on the other hand they must provide adequate safeguards to prevent practitioners' having their careers and livelihoods inappropriately disrupted. A clear understanding of the values inherent in the decision-making process may lead to more effective, efficient, and defensible approaches to recertification.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,016
score de la tête « metaresearch » (Gemma)0,085
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Évaluation · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Commentaire · Signal consensuel: aucune
Score de désaccord entre enseignants0,984
Score d'incertitude au seuil0,169

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0160,085
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0030,002
Études des sciences et des technologies0,0020,007
Communication savante0,0070,007
Science ouverte0,0020,005
Intégrité de la recherche0,0040,005
Charge utile insuffisante (le modèle a refusé de juger)0,0500,018

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,051
Tête enseignante GPT0,372
Écart entre enseignants0,321 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
DomaineÉvaluation
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations9
Publié2000
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revueAcademic MedicineMême sujetInnovations in Medical EducationTravaux en français237 207