How do validated measures of functional outcome compare with commonly used outcomes in administrative database research for lumbar spinal surgery?
Bibliographic record
Abstract
Clinical interpretation of health services research based on administrative databases is limited by the lack of patient-reported functional outcome measures. Reoperation, as a surrogate measure for poor outcome, may be biased by preferences of patients and surgeons and may even be planned a priori. Other available administrative data outcomes, such as postoperative cross sectional imaging (PCSI), may better reflect changes in functional outcome. The purpose was to determine if postoperative events captured from administrative databases, namely reoperation and PCSI, reflect outcomes as derived by validated functional outcome measures (short form 36 scores, Oswestry disability index) for patients who underwent discretionary surgery for specific degenerative conditions of the lumbar spine such as disc herniation, spinal stenosis, degenerative spondylolisthesis, and isthmic spondylolisthesis. After reviewing the records of all patients surgically treated for disc herniation, spinal stenosis, degenerative spondylolisthesis, and isthmic spondylolisthesis at our institution, we recorded the occurrence of PCSI (MRI or CT-myelograms) and reoperations, as well as demographic, surgical, and functional outcome data. We determined how early (within 6 months) and intermediate (within 18 months) term events (PCSI and reoperations) were associated with changes in intermediate (minimum 1 year) and late (minimum 2 years) term functional outcome, respectively. We further evaluated how early (6-12 months) and intermediate (12-24 months) term changes in functional outcome were associated with the subsequent occurrence of intermediate (12-24 months) and late (beyond 24 months) term adverse events, respectively. From 148 surgically treated patients, we found no significant relationship between the occurrence of PCSI or reoperation and subsequent changes in functional outcome at intermediate or late term. Similarly, earlier changes in functional outcome did not have any significant relationship with subsequent occurrences of adverse events at intermediate or late term. Although it may be tempting to consider administrative database outcome measures as proxies for poor functional outcome, we cannot conclude that a significant relationship exists between the occurrence of PCSI or reoperation and changes in functional outcome.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.478 | 0.706 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.006 |
| Bibliometrics | 0.014 | 0.023 |
| Science and technology studies | 0.002 | 0.009 |
| Scholarly communication | 0.015 | 0.013 |
| Open science | 0.004 | 0.005 |
| Research integrity | 0.005 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".