242. A systematic review and meta-analysis of models to predict the diagnosis of giant cell arteritis
Bibliographic record
Abstract
Objectives: The diagnosis of giant cell arteritis (GCA) can be difficult in individuals with inconclusive symptoms and inflammatory markers. Temporal artery biopsy (TAB), while often thought of as the gold standard, may miss a proportion of diagnoses. As such, many diagnostic models incorporating various predictors have been developed to assist clinicians in predicting a diagnosis of GCA with varying utility. This systematic review seeks to analyze these models to understand common and stable predictors of a diagnosis of GCA, understand their rigor, and determine how different criteria can alter the diagnosis of GCA. Methods: We performed a literature from January 1990 to May 2020 for studies that used a model to diagnose giant cell arteritis. Studies with models that had fewer than three variables or 30 people were excluded. Abstract screening, data extraction, and risk of bias were performed by two independent reviewers for each study. Study characteristics, patient characteristics, method of and criteria for diagnosis, and model details were extracted and summarized. Meta-analysis of individual signs and symptoms was performed using generic inverse variance. The PROBAST tool was used to assess risk of bias in each individual study. Results: We screened 1 446 abstracts and included 34 studies using data from 11 countries. 42 diagnostic models were identified. A total of 13 388 patients, 12 570 TABs, and 3 718 diagnoses of GCA were included. 22 studies required TAB positivity to diagnose GCA, 7 diagnosed using a composite of clinical and investigative findings, and 4 only required clinical findings. Rates of diagnosis of GCA were 25.0%, 39.0%, and 44.9% in each group respectively; Rates of TAB positive diagnoses was 98.2%, 53.7%, and 69.8%. There were 82.9% more diagnoses of GCA when using composite criteria over TAB positivity alone. Jaw claudication and Temporal changes were most associated with a diagnosis of GCA, however there were more predictive of TAB positive GCA than a clinical diagnoses, whereas headache and vision loss were more associated with non-TAB based diagnoses of GCA. 22 studies were at high risk of model bias and 4 were low risk. Conclusions: Models used to predict a diagnosis of GCA are of variable methodological quality and are largely dependent on using TAB positivity as a gold standard for a diagnosis of GCA. Despite this, predictors of GCA are consistent. Future models should focus on validation and use diagnostic standards that include composite criteria that reflect current practice. Disclosures: NK – Trial support from Roche, BMS, Sanofi, Abbvie. All others - None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.024 | 0.075 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.013 | 0.043 |
| Bibliometrics | 0.008 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".