Expanding the scope of quality measurement in surgery to include nonoperative care: Results from the American College of Surgeons National Surgical Quality Improvement Program emergency general surgery pilot
Bibliographic record
Abstract
BACKGROUND: Patients managed nonoperatively have been excluded from risk-adjusted benchmarking programs, including the American College of Surgeons (ACS) National Surgical Quality Improvement Program (NSQIP). Consequently, optimal performance evaluation is not possible for specialties like emergency general surgery (EGS) where nonoperative management is common. We developed a multi-institutional EGS clinical data registry within ACS NSQIP that includes patients managed nonoperatively to evaluate variability in nonoperative care across hospitals and identify gaps in performance assessment that occur when only operative cases are considered. METHODS: Using ACS NSQIP infrastructure and methodology, surgical consultations for acute appendicitis, acute cholecystitis, and small bowel obstruction (SBO) were sampled at 13 hospitals that volunteered to participate in the EGS clinical data registry. Standard NSQIP variables and 16 EGS-specific variables were abstracted with 30-day follow-up. To determine the influence of complications in nonoperative patients, rates of adverse outcomes were identified, and hospitals were ranked by performance with and then without including nonoperative cases. RESULTS: Two thousand ninety-one patients with EGS diagnoses were included, 46.6% with appendicitis, 24.3% with cholecystitis, and 29.1% with SBO. The overall rate of nonoperative management was 27.4%, 6.6% for appendicitis, 16.5% for cholecystitis, and 69.9% for SBO. Despite comprising only 27.4% of patients in the EGS pilot, nonoperative management accounted for 67.7% of deaths, 34.3% of serious morbidities, and 41.8% of hospital readmissions. After adjusting for patient characteristics and hospital diagnosis mix, addition of nonoperative management to hospital performance assessment resulted in 12 of 13 hospitals changing performance rank, with four hospitals changing by three or more positions. CONCLUSION: This study identifies a gap in performance evaluation when nonoperative patients are excluded from surgical quality assessment and demonstrates the feasibility of incorporating nonoperative care into existing surgical quality initiatives. Broadening the scope of hospital performance assessment to include nonoperative management creates an opportunity to improve the care of all surgical patients, not just those who have an operation. LEVEL OF EVIDENCE: Care management, level IV; Epidemiologic, level III.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".