External Validation of Scoring Instruments for Evaluating Pediatric Resuscitation
Notice bibliographique
Résumé
INTRODUCTION: Although many methods have been proposed to assess clinical performance during resuscitation, robust and generalizable metrics are still lacking. Further research is necessary to develop validated clinical performance assessment tools and show an improvement in outcomes after training. We aimed to establish evidence for validity of a previously published scoring instrument--the Clinical Performance Tool (CPT)--designed to evaluate clinical performance during simulated pediatric resuscitations. METHODS: This was a prospective experimental trial performed in the simulation laboratory of a pediatric tertiary care facility, with a pretest/posttest design that assessed residents before and after pediatric advanced life support (PALS) certification. Thirteen postgraduate year 1 (PGY1) and 11 PGY3 pediatric residents completed 5 simulated pediatric resuscitation scenarios each during 2 consecutive sessions; between the 2 sessions, they completed a full PALS certification course. All sessions were video recorded. Sessions were scored by raters using the CPT; total scores were expressed as a percentage of maximum points possible for each scenario. Validity evidence was established and interpreted according to Messick's framework. Evidence regarding relations to other variables was assessed by calculating differences in scores between pre-PALS and post-PALS certification and PGY1 and PGY3 using a repeated-measures analysis of variance test. Internal structure evidence was established by assessing interrater reliability using intraclass correlation coefficients (ICCs) for each scenario, a G-study, and a variance component analysis of individual measurement facets (scenarios, raters, and occasions) and associated interactions. RESULTS: Overall scores for the entire study cohort improved by 10% after PALS training. Scores improved by 9.9% (95% confidence interval [CI], 4.5-15.4) for the pulseless nonshockable arrest (ICC, 0.85; 95% CI, 0.74-0.92), 14.6% (95% CI, 6.7-22.4) for the pulseless shockable arrest (ICC, 0.98; 95% CI, 0.96-0.99), 4.1% (95% CI, -4.5 to 12.8) for the dysrhythmias (ICC, 0.92; 95% CI, 0.87-0.96), 18.4% (95% CI, 9.7-27.1) for the respiratory scenario (ICC, 0.97; 95% CI, 0.95-0.98), and 5.3% (95% CI, -1.4 to 2.0) for the shock scenarios (ICC, 0.94; 95% CI, 0.90-0.97). There were no differences between PGY1 and PGY3 scores before or after the PALS course. Reliability of the instrument was acceptable as demonstrated by a mean ICC of 0.95 (95% CI, 0.94-0.96). The G-study coefficient was 0.94. Most variance could be attributed to the subject (57%). Interactions between subject and scenario and subject and occasion were 9.9% and 1.4%, respectively, and variance attributable to rater was minimal (0%). CONCLUSIONS: Pediatric residents improved scores on CPT after completion of a PALS course. Clinical Performance Tool scores are sensitive to the increase in skills and knowledge resulting from such a course but not to learners' levels. Validity evidence from scores for the CPT confirms implementation in new contexts and partially supports internal structure. More evidence is required to further support internal structure and especially to support relations with other variables and consequence evidence. Additional modifications should be made to the CPT before considering its use for high-stakes certification such as PALS.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,120 | 0,163 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,003 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».