MétaCan
Menu
Back to cohort
Record W2159481738 · doi:10.1093/rheumatology/keq231

The revised BILAG Index with numerical scoring in systemic lupus erythematosus: added value with some limitations

2010· letter· en· W2159481738 on OpenAlexaff
Janet Pope

Bibliographic record

VenueLara D. Veeken · 2010
Typeletter
Languageen
FieldMedicine
TopicSystemic Lupus Erythematosus Research
Canadian institutionsSt Joseph's Health Care
Fundersnot available
KeywordsMedicineIndex (typography)Value (mathematics)Lupus erythematosusSystemic lupusSystemic lupus erythematosusInternal medicineImmunologyStatisticsAntibodyDisease

Abstract

fetched live from OpenAlex

The old index versus the BILAG-2004 and the fit of BILAG in SLE trials This editorial refers to ‘Numerical scoring for the BILAG-2004 index’, by Yee et al., doi:10.1093/rheumatology/keq026, on page 1665. Scoring disease activity in SLE is a laudable goal, as it is a complex disease potentially involving many organs. There have been quantitative scales such as the SLAM [1] and SLEDAI [2], and ordinal subsets of the BILAG [3], which has eight organ-based systems. In the original BILAG, there was no way to add individual organ scores together. The previous numerical scoring developed was A = 12, B = 5, C = 1, D = 0 and E = 0 [4]. In this issue of Rheumatology, Yee et al. [5] studied the numerical scoring of the BILAG-2004. Scoring using A = 12, B = 8, C = 1, and D and E = 0 gave the best fit, with the physician altering treatment as the main outcome. Data were derived from repeat and cross-sectional visits but did not consider change over time, as each visit was considered as distinct. One cannot determine how well the BILAG-2004 will function prospectively and if treatment will change when the score changes by four points (equivalent to one B going to an A) or seven points (changing from C to B). The analysis did not take into account previous disease activity, current treatment and patient preference of treatment. The minimum important numerical change has not been determined. Changes could be bi-directionally different and may vary depending on the number of baseline As and Bs, and total score. Undoubtedly, in clinical practice, the BILAG can determine when a change in therapy is needed [6], and it may also be better than the SLEDAI to detect the need to alter treatment [7]. As the BILAG validation was developed by a group that has been working with the BILAG for years, it is likely that results would be less optimal when others who are not well versed with BILAG use it. Moreover, the discontinuous nature of scoring A, B and C along with numerical weighting in a newer version may yield a problem: a score should be continuous, yet many parts of the BILAG are dichotomous (yes or no). The inability to give an overall score in the old BILAG can be overcome by the numerical scoring system of the BILAG-2004. However, problems still exist in the BILAG-2004 including interpretation of the variables. For some organs, an A is obvious as it is well described; however, in inflammatory arthritis an A is a loss of function. Moreover, as the next scenario shows, there could be various interpretations when scoring. A woman has lupus with Jaccoud’s arthropathy (subluxed MCPs); she is treated with an effective medication and her swollen joint count (SJC) reduces from 22/28 to 4/28. However, she is still functionally impaired due to subluxations. It is not totally certain whether her A has stayed an A (due to loss of function) or reduced to B. In other scales (SLAM and SLEDAI), her arthritis score would also remain unchanged. If the patient had RA and improved from 22/28 to 4/28, we would assume that the treatment was successful. The numerical BILAG-2004 also places value judgements on weighting of each organ scale. Some A scores are less severe than other organ scales. Weighting (although statistically determined) may lack face validity. Significant cytopenias or active proliferative GN may be more important as an A than inflammatory arthritis. Therefore, the variability of severity for different A scores can have an effect on the entry criteria into randomized controlled trials (RCTs) and responsiveness to treatments. Performing the BILAG is time consuming, requires training and the time frame is problematic when a patient may be in good health at her current visit but mentions that 2 weeks ago she had significant fever from SLE, which has now resolved. The changing nature of symptoms can affect the score, where an improving patient may not be scored as improving until a subsequent visit. Most clinicians will not use it in routine care. So, why use it at all one may ask? There is value in cohort studies and in clinical trials, but the BILAG may need large improvements to change the score. The BILAG (SLAM and SLEDAI) was not developed for use in trials but was used in cohort studies. The sensitivity to change is not high in some parts of each scale. For instance, if in an LN trial, baseline creatinine was twice the upper limit of normal with 4 g/day of proteinuria, treatment was given and the creatinine improved but did not return to a normal level and 1 g/day of proteinuria persisted. Prednisone dose was reduced by 75%, which would be considered a partial success by many. The time to fully resolve proteinuria is often beyond the duration of a trial. This point illustrates the problematic nature of lupus scales when used in RCTs in which the scores may not have changed despite obvious improvement. Not everyone agrees with the definitions within the BILAG. For instance, rash may include urticaria or not. RP is more apt to fluctuate seasonally than with changes in disease activity. In the past, lupus studies were negative and likely to be under-powered with insensitive outcome measures, but we need to move beyond this to avoid Type II errors [8]. So outcomes need to detect minimally important changes such as small-to-moderate improvements in the SJC or a modest reduction in daily proteinuria. The characteristics of the scale vary by organ system, and thus sample size calculations have to take into account what type of patients will be included and what improvement is reasonable to expect. The BILAG score was unable to differentiate active drug from placebo in some trials (which could be due to ineffectiveness of the drugs, protocols that minimized differences by steroid loading and later steroid withdrawal or due to lack of sensitivity of the BILAG to detect change). The BILAG-2004 is unlikely to overcome this limitation as it numerically replaces a score that was previously a letter, but does not detect modest differences in patients, which may be important in clinical trials. A recent trial showed a low placebo response with the BILAG despite steroid loading, which is unexpected [9]. There are flares, and then there are FLARES! Flares can be a problem as a BILAG A score suggests a change in therapy. In a rituximab trial, half the time a patient with a BILAG A had no change in treatment, whereas in clinical practice 92% of patients with flares and an A score had a change in treatment [10]. Thus, the BILAG is not accurate in determining change in treatment in clinical trials. Although recent equivalence trials have used the BILAG, we cannot fully establish whether the BILAG should have shown differences between groups. The BILAG was not different in patients receiving AZA compared with ciclosporin for steroid-sparing effects [11] or CYC vs MMF [12]. We have seen a trend to use the BILAG as a secondary outcome (i.e. no worsening in the BILAG score) and a steroid-sparing effect of treatment or improvement in physician global assessment and SLEDAI as a primary outcome. In this responder index, there can be no new BILAG As and no more than one new BILAG B [13]. Moreover, the heterogeneity of SLE plays a role in treatment response. It is unlikely that SLE of the skin will respond to the same therapies as renal disease. This problem increases the complexity in developing outcome measures for clinical trials. So, how do we move forward? The BILAG is a guide and common sense will prevail if there is discordance between clinical decision-making and the BILAG score. Other sites can consider validation of the numerical scoring of the BILAG-2004 in prospective studies. However, the BILAG is probably not an ideal primary outcome measure for SLE trials due to its inherent limitations. Disclosure statement: J.P. has participated in sponsored SLE trials with Teva, GSK and Roche.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.010
metaresearch head score (Gemma)0.054
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.010
Threshold uncertainty score0.051

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0100.054
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0020.002
Science and technology studies0.0000.001
Scholarly communication0.0020.002
Open science0.0020.001
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.028
GPT teacher head0.273
Teacher spread0.245 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2010
Admission routes1
Has abstractno

Explore more

Same venueLara D. VeekenSame topicSystemic Lupus Erythematosus ResearchFrench-language works237,207