The interrater reliability of static palpation of the thoracic spine for eliciting tenderness and stiffness to test for a manipulable lesion
Bibliographic record
Abstract
Despite widespread use by manual therapists, there is little evidence regarding the reliability of thoracic spine static palpation to test for a manipulable lesion using stiffness or tenderness as diagnostic markers. We aimed to determine the interrater agreement of thoracic spine static palpation for segmental tenderness and stiffness and determine the effect of standardised training for examiners. The secondary aim was to explore expert consensus on the level of segmental tenderness required to locate a “manipulable lesion”. Two experienced chiropractors used static palpation of thoracic vertebrae on two occasions (pragmatic and standardised approaches). Participants rated tenderness on an 11-point numerical pain rating scale (NPRS) and raters judged segmental stiffness based on their experience and perception of normal mobility with the requested outcomes of hypomobile or normal mobility. We calculated interrater agreement using percent agreement, Cohen’s Kappa coefficients ( κ ) and prevalence-adjusted bias-adjusted Kappa coefficients (PABAK). In a preliminary study, an expert panel of 10 chiropractors took part in a Delphi process to identify the level of meaningful segmental tenderness required to locate a “manipulable lesion”. Thirty-six participants (20 female) were enrolled for the reliability study on the 13th March 2017. Mean (SD) age was 22.4 (3.4) years with an equal distribution of asymptomatic ( n = 17) and symptomatic (n = 17) participants. Overall, the interrater agreement for spinal segmental stiffness had Kappa values indicating less than chance agreement [ κ range − 0.11, 0.53]. When adjusted for prevalence and bias, the PABAK ranged from slight to substantial agreement [0.12–0.76] with moderate or substantial agreement demonstrated at the majority of spinal levels (T1, T2 and T6 to T12). Generally, there was fair to substantial agreement for segmental tenderness [Kappa range 0.22–0.77]. Training did not significantly improve interrater agreement for stiffness or tenderness. The Delphi process indicated that an NPRS score of 2 out of 10 identified a potential “manipulable lesion”. Static palpation was overall moderately reliable for the identification of segmental thoracic spine stiffness and tenderness, with tenderness demonstrating a higher reliability. Also, an increased agreement was found within the mid-thoracic spine. A brief training intervention failed to improve reliability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".