A Machine Learning Framework for Intraoperative Segmentation and Quality Assessment of Pedicle Screw X-Rays
Bibliographic record
Abstract
Pedicle screw fixation is a technically demanding procedure with potential difficulties and reoperation rates are currently on the order of 11%. The most common intraoperative practice for position assessment of pedicle screws is biplanar fluoroscopic imaging that is limited to two-dimensions and is associated to low accuracies. We have previously introduced a full-dimensional position assessment framework based on registering intraoperative X-rays to preoperative volumetric images with sufficient accuracies. However, the framework requires a semi-manual process of pedicle screw segmentation and the intraoperative X-rays have to be taken from defined positions in space in order to avoid pedicle screws’ head occlusion. This motivated us to develop advancements to the system to achieve higher levels of automation in the hope of higher clinical feasibility. In this study, we developed an automatic segmentation and X-ray adequacy assessment protocol. An artificial neural network was trained on a dataset that included a number of digitally reconstructed radiographs representing pedicle screw projections from different points of view. This model was able to segment the projection of any pedicle screw given an X-ray as its input with accuracy of 93% of the pixels. Once the pedicle screw was segmented, a number of descriptive geometric features were extracted from the isolated blob. These segmented images were manually labels as ‘adequate’ or ‘not adequate’ depending on the visibility of the screw axis. The extracted features along with their corresponding labels were used to train a decision tree model that could classify each X-ray based on its adequacy with accuracies on the order of 95%. In conclusion, we presented here a robust, fast and automated pedicle screw segmentation process, combined with an accurate and automatic algorithm for classifying views of pedicle screws as adequate or not. These tools represent a useful step towards full automation of our pedicle screw positioning assessment system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".