Recognizing multiple needle insertion attempts for performance assessment in central venous catheterization training
Bibliographic record
Abstract
Purpose: Computer-assisted skill assessment has traditionally been focused on general metrics related to tool motion and usage time. While these metrics are important for an overall evaluation of skill, they do not address critical errors made during the procedure. This study examines the effectiveness of utilizing object detection to quantify the critical error of making multiple needle insertion attempts in central venous catheterization. Methods: 6860 images were annotated with ground truth bounding boxes around the syringe attached to the needle. The images were registered using the location of the phantom, and the bounding boxes from the training set were used to identify the regions where the needle was most likely inserting the phantom. A Faster region-based convolutional neural network was trained to identify the syringe and produce the bounding box location for images in the test set. A needle insertion attempt began when the location of the predicted bounding box fell within the identified insertion region. To evaluate this method, we compared the computed number of insertions to the number of insertions identified by human reviewers. Results: The object detection network had an overall mean average precision (mAP) of 0.71. This tracking method computed an average of 4.40 insertion attempts per recording compared to a reviewer count of 1.39 attempts per recording. Conclusions: The difference in the number of insertion attempts identified by the computer and reviewers decreases with an increasing mAP, making this method suitable for detecting multiple needle insertions using an object detection network with a high accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".