Measuring surgical outcomes in neurosurgery: implementation, analysis, and auditing a prospective series of more than 5000 procedures
Bibliographic record
Abstract
OBJECT: Health care reform debate includes discussions regarding outcomes of surgical interventions. Yet quality of medical care, when judged as a health outcome, is difficult to define because of impediments affecting accuracy in data collection, analysis, and reporting. In this prospective study, the authors report the outcomes for neurosurgical treatment based on point-of-care interactions recorded in the electronic medical record (EMR). METHODS: The authors' neurosurgery practice collected outcome data for 19 physicians and ancillary personnel using the EMR. Data were analyzed for 5361 consecutive surgical cases, either elective or emergency procedures, performed during 2009 at multiple hospitals, offices, and an ambulatory spine surgery center. Main outcomes included complications, length of stay (LOS), and discharge disposition for all patients and for certain frequently performed procedures. Physicians, nurses, and other medical staff used validated scales to record the hospital LOS, complications, disposition at discharge, and return to work. RESULTS: Of the 5361 surgical procedures performed, two-thirds were spinal procedures and one-third were cranial procedures. Organization-wide compliance with reporting rates of major complications improved throughout the year, from 80.7% in the first quarter to 90.3% in the fourth quarter. Auditing showed that rates of unreported complications decreased from 11% in the first quarter to 4% in the fourth quarter. Complication data were available for 4593 procedures (85.7%); of these, no complications were reported in 4367 (95.1%). Discharge dispositions reported were home in 86.2%, rehabilitation center in 8.9%, and nursing home in 2.5%. Major complications included culture-proven infection in 0.61%, CSF leak in 0.89%, reoperation within the same hospitalization in 0.38%, and new neurological deficits in 0.77%. For the commonly performed procedures, the median hospital LOS was 3 days for craniotomy for aneurysm or intraaxial tumor and less than 1 day for angiogram, anterior cervical discectomy with fusion, or lumbar discectomy. CONCLUSIONS: With prospectively collected outcome data for more than 5000 surgeries, the authors achieved their primary end point of institution-wide compliance and data accuracy. Components of this process included staged implementation with physician pilot studies and oversight, nurse participation, point-of-service data capture, EMR form modification, data auditing, and confidential surgeon reports.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".