ISSLS Prize Winner: Consensus on the Clinical Diagnosis of Lumbar Spinal Stenosis
Bibliographic record
Abstract
STUDY DESIGN: Delphi. OBJECTIVE: The aim of this study was to obtain an expert consensus on which history factors are most important in the clinical diagnosis of lumbar spinal stenosis (LSS). SUMMARY OF BACKGROUND DATA: LSS is a poorly defined clinical syndrome. Criteria for defining LSS are needed and should be informed by the experience of expert clinicians. METHODS: Phase 1 (Delphi Items): 20 members of the International Taskforce on the Diagnosis and Management of LSS confirmed a list of 14 history items. An online survey was developed that permits specialists to express the logical order in which they consider the items, and the level of certainty ascertained from the questions. Phase 2 (Delphi Study) Round 1: Survey distributed to members of the International Society for the Study of the Lumbar Spine. Round 2: Meeting of 9 members of Taskforce where consensus was reached on a final list of 10 items. Round 3: Final survey was distributed internationally. Phase 3: Final Taskforce consensus meeting. RESULTS: A total of 279 clinicians from 29 different countries, with a mean of 19 (±SD: 12) years in practice participated. The six top items were "leg or buttock pain while walking," "flex forward to relieve symptoms," "feel relief when using a shopping cart or bicycle," "motor or sensory disturbance while walking," "normal and symmetric foot pulses," "lower extremity weakness," and "low back pain." Significant change in certainty ceased after six questions at 80% (P < .05). CONCLUSION: This is the first study to reach an international consensus on the clinical diagnosis of LSS, and suggests that within six questions clinicians are 80% certain of diagnosis. We propose a consensus-based set of "seven history items" that can act as a pragmatic criterion for defining LSS in both clinical and research settings, which in the long term may lead to more cost-effective treatment, improved health care utilization, and enhanced patient outcomes. LEVEL OF EVIDENCE: 2.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".