International Medical Graduatesʼ Performances of Techniques of Physical Examination, with a Comparison of U.S. Citizens and Non—U.S. Citizens
Notice bibliographique
Résumé
Literature dating back over 25 years has documented and commented upon deficiencies in the performances of medical students and house officers in both techniques of physical examination and ability to detect abnormalities.1,2,3,4,5,6,7 Many of these studies took place in U.S. teaching hospitals, and when the subjects were residents, the studies do not specify results for international medical graduates (IMGs). Yet in 1998–1999 25% of first-year house officers in U.S. postgraduate medical training programs were IMGs,8 whose clinical experience in medical school is considered more variable than that offered in U.S. and Canadian schools.9 A substantial number of IMGs are actually American and Canadian citizens acquiring their undergraduate medical training outside North America, particularly at the “offshore” schools, which have proliferated in the Caribbean. The quality of training in clinical skills in these new schools, which are not accredited by the Liaison Council on Medical Education, is largely unknown. In July 1998 the Educational Commission for Foreign Medical Graduates (ECFMG) implemented its Clinical Skills Assessment (CSA®), a new requirement for ECFMG certification. Experience with a more recently created “physical examination case” within the CSA has allowed us to measure skills in a selection of basic physical examination techniques among IMGs completing this high-stakes performance assessment, including non-U.S. citizens and U.S. citizens. Both groups were deficient in important skills. Methods Test Case and Design. The CSA is a ten-station performance assessment using standardized patients (SPs). It is designed to measure capabilities in history taking, certain aspects of physical examination, oral and written communications, interpersonal behaviors, and the English language. The typical case requires the examinee to assess a new patient problem by taking a focused history and performing what the examinee considers a relevant physical examination. The examinee completes a “patient note” and suggests a differential diagnosis and a diagnostic plan. The SPs use checklists to document which expected elements of history taking and physical diagnosis the candidate did, but the format cannot always distinguish whether an examinee omitted a physical examination element or attempted it but made an error. For this reason, the physician staff of the CSA, with the endorsement of its Test Development Committee, created a “physical examination case.” The case scenario presents a young man who needs a pre-employment physical examination; the “patient” hands to the examinee a simulated examination form that explicitly indicates the physical examination components to be done (listed in Table 1). These elements were chosen not to replicate an entirely realistic pre-employment examination, but rather to include tasks relating to a variety of organ systems and to include some for which correct technique would likely be especially necessary for detecting abnormalities in actual practice.TABLE 1: Performances of Non-U.S.-Citizen and U.S.-Citizen International Medical Graduates (IMGs) of Selected Physical Examination Techniques Tested in One Case within the Clinical Skills Assessment (CSA®) of the Educational Commission for Foreign Medical Graduates, October 1999–January 2000Each task, such as “auscultation of lungs,” “ophthalmoscopic examination,” “deep tendon reflexes,” was broken down into one to four components or scoring criteria, each scored by the SP as “done” or “not done” by the examinee. For example, for ophthalmoscopic examination, a candidate can separately obtain a point for correctly instructing the patient, for using his or her right eye for patient's right eye and left for left, and for bringing the instrument sufficiently close to the patient's eye. Such criteria were largely based on techniques outlined in a standard textbook on physical examination that is used extensively both in the United States and in other countries.10 Criteria for auscultation of the heart were minimal because we could not fairly expect special positioning and maneuvers in the setting of a screening examination for a young adult without complaints. More criteria might have been entered for some tasks, but since the CSA intends to assess basic clinical skills, we aimed at the most essential techniques. Also, we did not want to impair SP recall with too long a checklist. The case was carried out by one SP following intensive training. One author (SJP) validated the SP's accuracy using simultaneous checklist scoring during “pilot” runs of the case. The SP was already considered by staff a rapid learner and accurate in his work, and showed good scoring concurrence during quality-assurance observations in this and another case (less than 10% discrepancy). He has no physical abnormalities. The case is not used in every administration of the test. It is one of a group of “miscellaneous” cases chosen by a computerized selection program designed to achieve balanced forms while accommodating the availability of SPs. From the perspective of any one candidate, the appearance of this case on her or his ten-case form was effectively random: the candidates were in no way prospectively selected. We report on the first 318 candidates who encountered this case on their form—247 non-U.S. citizens and 71 U.S. citizens from October 1999 through January 2000. The ratio of U.S. to non-U.S. citizens (.29) in this group turned out to be somewhat lower than the ratio (.44) for all 8,313 candidates tested by the CSA as of the date of analysis. The overall test scores in data gathering (history taking and physical examination) across all ten cases in their forms for the 318 examinees in this study were similar to those for all candidates (t1,8632 = 1,732, p = 0.08), suggesting that our cohort was representative. Analysis. For a task, such as “deep tendon reflexes,” comprising four scoring criteria, an examinee could obtain 0, 1, 2, 3, or 4 points, expressed for each task as a percent-correct score of 0%, 25%, 50%, 75%, or 100%, with a similar transformation used for tasks comprising fewer subtasks. We calculated the mean of these percent-correct scores for each task (e.g., deep tendon reflexes), and for the whole case (percentage of all 19 criteria done correctly), over all examinees in the cohort, and did the same for the sub-groups of U.S. IMGs and non-U.S. IMGs. Confidence intervals were also calculated. To compare the performances of non-U.S. IMGs and U.S. IMGs, we conducted a repeated-measures analysis of variance. The eight physical examination tasks were the within-subject factors, and citizenship at start of medical school was the between-subjects factor. Post-hoc analyses were conducted to determine whether differences in task scores between groups were significant. To better understand qualitatively the nature of frequently scored errors and omissions, the author (SJP) most responsible for designing the case and training the SP interviewed the SP and observed 40 randomly selected tapes of actual encounters. The SP was asked, where appropriate, to recall the most common errors causing him to withhold a mark for a given scoring criterion (e.g., “palpating too high on the foot” for dorsalis pedis pulse). Results Table 1 shows the mean percentage scores for each task and for the whole case for all examinees in the cohort and for the U.S. and non-U.S. subgroups. The task main effect was significant (F = 22.631, p <.01), indicating that the tasks, averaged over the two groups, were not of equal difficulty. The weakest performance was in ophthalmoscopy, the strongest in cardiac examination (for which, as mentioned, the criteria were minimal). There was a significant group (between-subjects) effect (F = 14.325, p <.01), indicating that, averaged over the eight tasks (or the whole case), there was a statistically significant difference in scores between the two groups. The U.S. IMGs obtained significantly higher case scores than did the non-U.S. IMGs. The group-by-task interaction was also significant (F = 4.126, p <.01), indicating that differences in performances between groups varied over the eight tasks. That is, the U.S. IMGs performed significantly better than did the non-U.S. IMGs for extraocular movements, ophthalmoscopic examination, locating radial and dorsalis pedis pulses, and deep tendon reflexes. Analysis of the scores for each scoring criterion within the eight physical examination tasks (not presented here), a “debriefing” interview with the SP, and review of a sample of videotapes of encounters revealed the following common technical deficiencies: clumsiness in properly placing and wrapping the blood pressure cuff; insufficient extent of induction of motion in testing eye movement; failure to use “right eye for right eye and left eye for left eye” and not bringing the instrument in closely enough for ophthalmoscopic examination; not comparing right with left at a given location on the thorax for pulmonary percussion; unfamiliarity with the location of the dorsalis pedis pulse; lack of briskness in applying the reflex hammer and applications at incorrect locations. Discussion Our study looked only at proficiencies in some fundamental techniques of physical examination, not the ability to recognize and interpret abnormalities. Thus, performance levels below 90% of criteria met can be considered a cause for some alarm when observed in medical school graduates, or final-year students, intending to enter a postgraduate training program.6 While few prescribed and traditional techniques in physical diagnosis have been rigorously tested to determine whether they improve accuracy in detecting or excluding abnormalities, we used as criteria well-established methods advocated in the most widely used textbook of physical diagnosis. Furthermore, it is difficult to deny that little will be seen in the fundus by an examiner holding the instrument 10 inches from the eye, or that a meaningful interpretation of the deep tendon reflexes is unlikely to follow misapplication of the hammer. It is not too much to expect that every new house officer on the first day of residency would be able to effortlessly and rapidly apply and use the sphygmomanometer in an urgent situation, yet our cohort of IMGs showed only an 87% level of proficiency in this skill. Of interest, McKay et al. tested Canadian medical graduates and found deficiencies in the technique of blood pressure measurement, though they used a more stringent set of criteria than ours.11 The ophthalmoscopic examination warrants comment. Non-U.S. IMGs showed only a 60% and U.S. IMGs an 80% level of proficiency, a significant difference but low score for both. Recent literature5,7 and the observations of one of us (SJP) at the medical school where he teaches suggest a declining use of the ophthalmoscope among learners and teachers in American academic medicine. Our results in this study hint that the situation is similar elsewhere. McNaught and Pearson, in the United Kingdom, found that ownership of an ophthalmoscope declined sharply after an “equipment grant” was discontinued.12 While any conception of the core skills in physical diagnosis must evolve to match changing patterns of practice,7 arguably all general physicians and some non-ophthalmologic specialists should be able to recognize at least papilledema, the advanced optic cupping of glaucoma, and perhaps some of the findings associated with common diseases such as diabetes and hypertension. Faulty basic technique, as evidenced by the IMGs we tested, will both frustrate those trying to master this difficult element of physical diagnosis and impede accuracy. Why might non-U.S. IMGs have performed less well in some tasks than U.S. IMGs? The ECFMG elected to create and implement the CSA based in part on the belief that clinical instruction among international medical schools is less standardized and more variable in extent than that offered by U.S. and Canadian schools accredited by the Liaison Committee for Medical Education.9 A majority of U.S. IMGs taking the CSA have attended one of the “offshore” medical schools. Students in these schools do much of their third- and fourth-year clinical rotations in U.S. hospitals and practices, and so may encounter the sorts of physical diagnosis expectations tested for in the CSA. Candidates have the opportunity to try out the physical examination equipment available in our examination rooms before the examination begins. Staff have on several occasions heard non-U.S. IMGs report that they had never used an ophthalmoscope or (more rarely) had seldom performed a blood pressure measurement. We are not aware, however, of any comparison of preliminary clinical skills instruction among U.S./Canadian, “offshore,” and other international medical schools. We do not believe that the SP performing this case showed bias in favor of U.S. IMGs over non-U.S. candidates. Obviously our training program for SPs includes discussion of bias and the imperative to avoid it. Also, by chance, the SP chosen for this case is himself a native of another country and speaks with an accent. Furthermore, our observations of a sample of encounters seemed to confirm the differences detected. Our study has limitations. It provides no comparison of skills of IMGs with those of graduates of U.S. and Canadian schools, and we by no means intend to imply that the latter would not show some deficiencies—indeed, literature cited earlier suggests that they would. We were not able to assess all commonly used physical examination tasks, and such skills as rectal, pelvic, and breast examination are not incorporated into the CSA. As noted, our physical examination case yields little information on ability to carry out a thorough cardiac examination appropriate to a symptomatic patient. Observations of videotapes revealed that occasional candidates did not attend to the explicit instructions for the case and failed to attempt one or more tasks, though we do not think the resulting invalid scores would influence the overall results and conclusions. This study has several implications. First, residency program directors should be aware that some medical graduates entering their programs might not bring with them a full repertoire of fundamental skills in physical examination technique; of course, our results apply only to graduates of medical schools outside the United States and Canada. It therefore may be desirable to assess selected clinical skills early in the first year and provide focused remediation for detected errors. Second, those responsible for clinical skills instruction at the medical school level may need to also sharpen their focus on ensuring the acquisition of fundamental physical diagnosis methods before students graduate. Finally, the authors' experience with this station supports the now widely accepted view that well-trained standardized patients can be used to assess ability in at least rudimentary techniques of physical examination.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».