Comparison of accuracy, revision, and perioperative outcomes in robot-assisted spine surgeries: systematic review and meta-analysis
Bibliographic record
Abstract
OBJECTIVE: Pedicle screw placement guidance is critical in spinal fusions, and spinal surgery robots aim to improve accuracy and reduce complications. Current literature has yet to compare the relative merits of available robotic systems. In this review, the authors aimed to 1) assess the current state of spinal robotics literature; 2) conduct a meta-analysis of robotic performance based on accuracy, speed, and safety; and 3) offer recommendations for robotic system selection. METHODS: Following PRISMA guidelines, the authors conducted a systematic literature review across PubMed, Embase, Cochrane Library, Web of Science, and Scopus as of April 28, 2022, for studies on approved robots for placing lumbar pedicle screws. Three reviewers screened and extracted data relating to the study characteristics, accuracy rate, intraoperative revisions, and reoperations. Secondary performance metrics included operative time, blood loss, and radiation exposure. The authors statistically compared the performance of the robots using a random-effects model to account for variation within and between the studies. Each robot was also compared with performance benchmarks of traditional techniques including freehand, fluoroscopic, and CT-navigated insertion. Finally, we performed a Duval and Tweedie trim-and-fill test to assess for the presence of publication bias. RESULTS: The authors identified 46 studies, describing 4670 patients and 25,054 screws, that evaluated 4 different robotic systems: Mazor X, ROSA, ExcelsiusGPS, and Cirq. The weighted accuracy rates of Gertzbein-Robbins classification grade A or B screws were as follows: ExcelsiusGPS, 98.0%; ROSA, 98.0%; Mazor, 98.2%; and Cirq, 94.2%. No robot was significantly more accurate than the others. However, the accuracy of the ExcelsiusGPS was significantly higher than that of traditional methods, and the accuracies of the Mazor and ROSA were significantly higher than that of fluoroscopy. The intraoperative revision rates were Cirq, 0.55%; ROSA, 0.91%; Mazor, 0.98%; and ExcelsiusGPS, 1.08%. The reoperation rates were Cirq, 0.28%; ExcelsiusGPS, 0.32%; and Mazor, 0.76% (no reoperations were reported for ROSA). Operative times were similar for all robots. Both the ExcelsiusGPS and Mazor were associated with significantly less blood loss than the ROSA. The Cirq had the lowest radiation exposure. Robots tended to be more accurate and generally their use was associated with fewer reoperations and less blood loss than freehand, fluoroscopic, or CT-navigated techniques. CONCLUSIONS: Robotic platforms perform comparably based on key metrics, with high accuracy rates and low intraoperative revision and reoperation rates. The spinal robotics publication rate will continue to accelerate, and choosing a robot will depend on the context of the practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.022 | 0.004 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".