Referral Practices for Spinal Surgery are Poorly Predicted by Clinical Guidelines and Opinions of Primary Care Physicians
Bibliographic record
Abstract
BACKGROUND: Degenerative disease of the lumbar spine is common. Although surgery can benefit selected patients, variation in surgical referrals reduces overall access to care. OBJECTIVES: To compare the actual referral practices for patients with degenerative disease of the lumbar spine with recommendations for surgical referral based on clinical practice guidelines (CPGs) and family physician (FP) opinions. RESEARCH DESIGN: An expert panel of primary and specialist physicians, using a Delphi process, came to a consensus on referral recommendations from CPGs based on a series of clinical vignettes. The vignettes were also presented to practicing FPs in Ontario, Canada, to determine their preferences for (or likelihood of) referral. SUBJECTS: We assembled a 10-member multispecialty expert panel. Practicing FPs were randomly sampled, stratified by county, and their patients were sampled purposefully by the FP. MEASURES: Respondents, both panelists and FPs, were asked to rate the appropriateness of surgical referral for a series of clinical vignettes. Patients reported their clinical symptoms and whether they had been referred to a surgeon. Using random-effects probit regression, predictions were compared with actual referral. Receiver operating characteristic curves were constructed and area under the curve (AUC) was measured. RESULTS: Consensus of the panel on recommendations for referral was achieved after 2 iterations (Cronbach alpha = 0.96). Based on responses from 107 patients and 61 FPs, we found poor concordance of both predicted FP preferences (AUC 0.57) and CPG recommendations (AUC 0.64) with actual referral. CONCLUSIONS: Referral practices are poorly predicted by CPG recommendations and individual FP opinions, based on clinical factors. Understanding other nonclinical factors may be more important in reducing variation in referrals and improving access.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.067 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".