Rectal Dissection Simulator for da Vinci Surgery: Details of Simulator Manufacturing With Evidence of Construct, Face, and Content Validity
Bibliographic record
Abstract
BACKGROUND: Apprenticeship in training new surgical skills is problematic, because it involves human subjects. To date there are limited inanimate trainers for rectal surgery. OBJECTIVE: The purpose of this article is to present manufacturing details accompanied by evidence of construct, face, and content validity for a robotic rectal dissection simulation. DESIGN: Residents versus experts were recruited and tested on performing simulated total mesorectal excision. Time for each dissection was recorded. Effectiveness of retraction to achieve adequate exposure was scored on a dichotomous yes-or-no scale. Number of critical errors was counted. Dissection quality was tested using a visual 7-point Likert scale. The times and scores were then compared to assess construct validity. Two scorer results were used to show interobserver agreement. A 5-point Likert scale questionnaire was administered to each participant inquiring about basic demographics, surgical experience, and opinion of the simulator. Survey data relevant to the determination of face validity (realism and ease of use) and content validity (appropriateness and usefulness) were then analyzed. SETTINGS: The study was conducted at a single teaching institution. SUBJECTS: Residents and trained surgeons were included. INTERVENTION: The study intervention included total mesorectal excision on an inanimate model. MAIN OUTCOME MEASURES: Metrics confirming or refuting that the model can distinguish between novices and experts were measured. RESULTS: A total of 19 residents and 9 experts were recruited. The residents versus experts comparison featured average completion times of 31.3 versus 10.3 minutes, percentage achieving adequate exposure of 5.3% versus 88.9%, number of errors of 31.9 versus 3.9, and dissection quality scores of 1.8 versus 5.2. Interobserver correlations of R = 0.977 or better confirmed interobserver agreement. Overall average scores were 4.2 of 5.0 for face validation and 4.5 of 5.0 for content validation. LIMITATIONS: The use of a da Vinci microblade instead of hook electrocautery was a study limitation. CONCLUSIONS: The pelvic model showed evidence of construct validity, because all of the measured performance indicators accurately differentiated the 2 groups studied. Furthermore, study participants provided evidence for the simulator's face and content validity. These results justify proceeding to the next stage of validation, which consists of evaluating predictive and concurrent validity. See Video Abstract at http://links.lww.com/DCR/A551.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".