A Tool to Assess Competence in Critical Care Ultrasound Based on Entrustable Professional Activities
Bibliographic record
Abstract
Abstract Background Existing assessment tools for competence in critical care ultrasound (CCUS) have limited scope and interrupt clinical workflow. The framework of entrustable professional activities (EPAs) is well suited to developing an assessment tool that is comprehensive and readily integrated into the intensive care unit (ICU) training environment. Objective This study sought to design an EPA-based tool to assess competence in CCUS for pulmonary and critical care fellows and to assess the validity and reliability of the tool. Methods Eight experts in CCUS met to define the core EPAs for CCUS. A nominal group technique was used to reach consensus. An assessment tool was created based on the EPAs with a modified Ottawa entrustability scale. Trained faculty evaluated pulmonary and critical care fellows using this tool in the ICU over a 6-month study period at a single institution. An assessment of validity of the EPA-based tool is made with four sources of validity evidence: content, response process, reliability, and relation to other variables. Reliability and response process data were generated using generalizability theory analysis to estimate sources of variance in entrustment scores. Analysis of response process validity and validity by relation to other variables was performed using regression models. Results Fifty-four assessments were recorded during the study period, conducted on 23 trainees by 13 faculty. Content validity of the tool was demonstrated using expert consensus and published guidelines from critical care societies to define the EPAs. Response process validity was demonstrated by the low variance in entrustment scores due to evaluators (0.086 or 6%) and high agreement between score and trainee self-assessment (regression coefficient, 0.82; P < 0.0001). Reliability was demonstrated by the high “true” variance in entrustment score attributable to the trainee: 0.674 or 45%. Validity by relation to other variables was demonstrated using regression analysis to show correlation between entrustment score and the number of times a fellow has performed an EPA (regression coefficient, 0.023; P < 0.0001). Conclusion An EPA-based assessment tool for competence in CCUS was created. We obtained sufficient validity evidence on three of the diagnostic EPAs. Procedural EPAs were infrequently assessed, limiting generalizability in this subgroup.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".