Sex Bias in Diagnostic Delay: Are Axial Spondyloarthritis and Ankylosing Spondylitis Still Phantom Diseases in Women? A Systematic Review and Meta-Analysis
Bibliographic record
Abstract
Diagnostic delay (DD) is associated with poor radiological and quality of life outcomes in axial spondyloarthritis (ax-SpA) and ankylosing spondylitis (AS). The female (F) population is often misdiagnosed, as classification criteria were previously studied mostly in males (M). We conducted a systematic review to investigate (i) the difference in DD between the sexes, the impact of HLA*B27 and clinical and social factors (work and education) on this gap, and (ii) the possible influence of the year of publication (before and after the 2009 ASAS classification criteria), geographical region (Europe and Israel vs. extra-European countries), sample sources (mono-center vs. multi-center studies), and world bank (WB) economic class on DD in both sexes. We searched, in PubMed and Embase, studies that reported the mean or median DD or the statistical difference in DD between sexes, adding a manual search. Starting from 399 publications, we selected 26 studies (17 from PubMed and Embase, 9 from manual search) that were successively evaluated with the modified Newcastle–Ottawa Scale (m-NOS). The mean DD of 16 high-quality (m-NOS > 4/8) studies, pooled with random-effects meta-analysis, produces results higher in F (1.48, 95% CI 0.83–2.14, p < 0.0001) but with significant results at the second analysis only in articles published before the 2009 ASAS classification criteria (0.95, 95% CI 0.05–1.85, p < 0.0001) and in extra-European countries (3.16, 95% CI 2.11–4.22, p < 0.05). With limited evidence, some studies suggest that DD in F might be positively influenced by HLA*B27 positivity, peripheral involvement, and social factors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.021 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".