Towards Robust Validity Evidence for Learning Environment Assessment Tools
Notice bibliographique
Résumé
To the Editor: Colbert-Getz and colleagues’1 review of learning environment (LE) assessment tools is both timely and relevant. The authors use four of the five categories of validity evidence from the American Psychological and Educational Research Associations to arrive at a total validity evidence score for each tool they reviewed to judge the quality of both medical student and resident LE perceptions. We wonder, however, if the authors used only the initial publications to arrive at their total validity evidence scores. For example, while the Medical Student Learning Environment Survey’s (MSLES) original publication received a total score of 3/8 (38%), subsequent publications from Australia and Canada examined the Internal Structure and Relationship to Other Variables criteria.2,3 Both studies used factor analysis to independently confirm that the individual scales of the MSLES represent one dominant factor. Clarke et al2 also examined the retest reliability and internal consistency, while Rusticus et al3 correlated the MSLES to student satisfaction and academic performance. Applying the authors’ validity criteria, we would have given an additional rating of 2 (“strong” evidence) for Internal Structure and a score of 1 (“weak” evidence) for Relationship to Other Variables. The total validity evidence score of the MSLES would therefore increase to 6/8 (75%). Similarly, for measuring the resident LE, the authors give the VA Learners’ Perception Survey (LPS) a total validity score of 2/8 (25%). The original publication by Keitz et al4 used focus groups of medical students and residents in the initial development of the LPS, and factor analysis was used to collapse the original 57 questions into four major domains. Internal consistency using a mixed-effects model was further verified in a subsequent publication by Cannon et al.5 We would have given an additional rating of 1 for Response Process and 2 for Internal Structure, increasing the total validity evidence score to 5/8 (63%). Both of these updated scores for the MSLES and LPS would be the highest scores listed for validity evidence in undergraduate and graduate medical education, respectively. We also wonder if the authors assessed the interrater reliability used to assess the validity evidence, since their checklist was adapted from Beckman et al,6 which found kappa values ranging from −0.10 to 0.96 and was particularly poor for rating the Response Process criteria. Lawrence K. Loo, MD Vice chair, Education and Faculty Development, Department of Medicine, and professor of medicine, Loma Linda University School of Medicine, Loma Linda, California; [email protected] John M. Byrne, DO Associate chief of staff, Education, VA Loma Linda Healthcare System, and associate professor of medicine, Loma Linda University School of Medicine, Loma Linda, California.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,028 | 0,043 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,002 | 0,007 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».