MP66-19 INTER-OBSERVER RELIABILITY OF EXAMINER SCORING ON HIGH STAKES UROLOGY OBJECTIVE STRUCTURED CLINICAL EXAMINATION
Notice bibliographique
Résumé
You have accessJournal of UrologyCME1 Apr 2023MP66-19 INTER-OBSERVER RELIABILITY OF EXAMINER SCORING ON HIGH STAKES UROLOGY OBJECTIVE STRUCTURED CLINICAL EXAMINATION Charles Paco, Iain Macintyre, and Naji Touma Charles PacoCharles Paco More articles by this author , Iain MacintyreIain Macintyre More articles by this author , and Naji ToumaNaji Touma More articles by this author View All Author Informationhttps://doi.org/10.1097/JU.0000000000003329.19AboutPDF ToolsAdd to favoritesDownload CitationsTrack CitationsPermissionsReprints ShareFacebookLinked InTwitterEmail Abstract INTRODUCTION AND OBJECTIVE: The Objective Structured Clinical Examination (OSCE) is an attractive tool of competency assessment in a high stakes summative exam. An advantage of the OSCE is the ability to assess more realistic context, content and procedures. Each year, the Queen’s Urology Exam Skills Training (QUEST) is attended by graduating Canadian urology residents to simulate their upcoming board exams. The exam consists of a written component and an OSCE. The aim of this study was to determine the inter-observer consistency of scoring between two examiners of an OSCE for a given candidate. METHODS: 39 participants in 2020 and 37 participants in 2021 completed four stations OSCEs virtually over the Zoom platform. Each candidate was examined and scored independently by 2 different faculty urologists in a blinded fashion at each station. The OSCE scoring consisted of a checklist rating scale for each question. An intra class correlation (ICC) analysis was conducted to determine the inter-rater reliability of the two examiners for each of the four OSCE stations in both the 2020 and 2021 OSCEs. RESULTS: For the 2020 data, the prostate cancer station scores were most strongly correlated (ICC 0.746, 95% CI (0.556-0.862) p<0.001). This was followed by the general urology station (ICC 0.688, 95% CI (0.464-0.829) p<0.001, the urinary incontinence station (ICC 0.638 95% CI (0.403- 0.794) p<0.001) and finally the nephrolithiasis station 0.472 95% CI (0.183-0.686) p<0.001). For the 2021 data, the renal cancer station had the highest ICC at 0.866 (95% CI (0.754-0.930) p<0.001). This was followed by the nephrolithiasis station (ICC 0.817 95% CI (0.673-0.901) p<0.001), the pediatric station (ICC 0.809, 95% CI (0.660-0.897) p<0.001) and finally the andrology station (ICC 0.804, 95% CI (649-0.895) p<0.001). Values less than 0.5 are indicative of poor reliability, values between 0.5 and 0.75 indicate moderate reliability, values between 0.75 and 0.9 indicate good reliability, and values greater than 0.90 indicate excellent reliability. CONCLUSIONS: Given a specific clinical scenario in an OSCE exam, inter-rater reliability of scoring can be compromised on occasion. The factors determining this divergence in examiner agreement will need further research to help elucidate, especially if OSCEs are continued as the gold standard in high stakes examinations. Source of Funding: None © 2023 by American Urological Association Education and Research, Inc.FiguresReferencesRelatedDetails Volume 209Issue Supplement 4April 2023Page: e941 Advertisement Copyright & Permissions© 2023 by American Urological Association Education and Research, Inc.MetricsAuthor Information Charles Paco More articles by this author Iain Macintyre More articles by this author Naji Touma More articles by this author Expand All Advertisement PDF downloadLoading ...
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,019 | 0,056 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,000 | 0,002 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,011 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».