Professional Insights from a Pioneer in Autoimmune Disease Testing: The Future of Antinuclear/Anticellular Antibody Testing
Bibliographic record
Abstract
Almost half a century has lapsed since I embarked on a career-long study of systemic autoimmune rheumatic diseases (SARDs)2 with a focus on their autoantibodies directed against an astounding spectrum of cellular antigens (1). The discovery of the lupus erythematosus cell and the development of the lupus erythematosus cell test serves as a historic reference point for the study of antinuclear antibodies (ANAs), or what today international consensus advocated should more correctly be referred to as anticellular antibodies (ACAs) (2, 3). Paralleling the explosion of the spectrum of ACAs was a remarkable transition in the technologies used to detect autoantibodies (1). Although some of the “octogenarian” assays such as double immunodiffusion, hemagglutination, complement fixation, and counterimmunoelectrophoresis are fading into oblivion, the ACA indirect immunofluorescence (IIF) test is increasingly used as a screening test and entry criterion for SARD. However, the emergence of newer multianalyte array technologies that have higher throughput, sensitivity, and specificity and detect a broader range of autoantibodies in comparatively miniscule serum samples may eventually replace the ACA IIF test (4, 5). ACA testing was once regarded the primary domain of rheumatologists and clinical immunologists, but with a growing spectrum of autoimmune diseases linked to ACA and other biomarkers, the spectrum of clinicians using these tests has also remarkably widened (1). More than 180 target autoantibodies have been described in systemic lupus erythematosus and >30 in systemic sclerosis. Nevertheless, fewer than 15 ACA in each of these diseases are used in routine diagnostic assays. Some of these ACA have perished in the “death valley” of translation because they fail to meet criteria for clinical applications (4). The continuously expanding spectrum of ACAs in SARD might be considered by some as unnecessary, but, for one thing, these efforts are closing the “seronegative gap” (6). In addition, some of these “orphan” ACA may be rediscovered and find an important niche when artificial intelligence is applied to diagnostics in the future (6). As witnessed by recommendations endorsed by the American College of Rheumatology and the Canadian Rheumatology Association in the Choosing Wisely paradigm, and “consensus” statements of the American College of Physicians and the College of American Pathologists, ACA testing has been under considerable review and criticism (reviewed in 7, 8). Most recently, the Effective Health Care Program characterized the evidence about ACA testing as “broad clinical consensus that ANA testing (including ANA subserologies) should not be used to screen for SARDs in primary care”; therefore, “there is no clinical uncertainty that a new systematic review could potentially address” (8). In other words, despite evidence to the contrary (reviewed in 7), ACA testing should be curtailed (particularly in primary care) because of its “poor positive and negative predictive values (PPV, 29%; NPV, 77%), leading to increased healthcare costs with unclear clinical benefit.” With these issues in mind, my personal perspectives on the future of ACA testing is summarized in 4 questions. First, what should we do with evidence that ACA and related biomarkers antedate the diagnosis of SARD by up to 20 years? Unfortunately, the proclamations of Choosing Wisely and Effective Health Care arise from a rather myopic perspective that ANA testing should be done on only high-PPV/low-NPV SARD patients. Second, the circular logic is difficult to rationalize because, if the patient has high-PPV SARD, why should the ACA test be done at all? An important aspect that seems to be overlooked is that when the “intent to treat” approach to ACA testing is restricted to high-PPV SARD, the diagnosis of SARDs is often delayed so that a considerable proportion of patients have active disease and end organ damage at the time of diagnosis (reviewed in 9). This delay in diagnosis is associated with remarkably high healthcare costs because renal disease, pulmonary fibrosis and hypertension, and irreversible joint damage, to name a few, require much more intensive and expensive care attended by decreased health-related quality of life. These observations are prompting many clinicians to reconsider the approach to SARDs by making concerted efforts to make a much earlier diagnosis (6). It needs to be appreciated that an earlier diagnosis is the domain of primary healthcare providers who serve as the SARD “case finders” (7). Screening tests such as ACA for SARD are used as part of “case finding,” and then, on the basis of clinical acumen, patients are referred to subspecialists for evaluation and appropriate management. Third, if primary care physicians aren't the early SARD case finders in the real world where there is a severe shortage of tertiary care specialists, who is? And fourth, given the documented and perceived limitations of the ANA/ACA IIF test as a screen for SARD (10), what should replace it? Perceived abuse of ANA/ACA testing is leading to revised laboratory approaches to screen for SARDs. ANA/ACA testing needs to move beyond the paradigm of confirming a diagnosis in high-pretest probability patients and “intent to treat” to “case finding” of very early autoinflammatory disease in which the clinical paradigm is “intent to prevent” morbidity and thereby decrease healthcare costs. Practical and realistic considerations indicate a continuing key role for primary care health providers as “case finders” who then refer patients for further investigation and treatment. There are advantages and trends indicating that ANA/ACA IIF as a screening testing will be replaced by high-throughput multianalyte array technologies. As an abbreviated reply to the last question, modern laboratories are migrating to technology platforms that have higher throughput with faster turnaround times. While newer ANA/ACA IIF assay platforms have also moved in this direction (1), it will be a challenge to meet the performance and clinical value of newer multi-analyte array technologies (MAAT) (4). Recent evidence indicates that the best approach for ACA testing is to combine ACA by IIF with MAAT (1). In conclusion, there is a strong need for unbiased approaches to ACA testing, based on most recent evidence. The laboratory approaches of the future need to consider the importance of disease prevention fostered by “case finding” and attenuation of significant morbidity and healthcare expenditures. As MAAT improve and decrease in price, it is likely that ACA IIF will no longer be the SARD screening assay of choice. systemic autoimmune rheumatic diseases antinuclear antibodies anticellular antibodies indirect immunofluorescence positive predictive value negative predictive value. The author's career and any measure of success would not have been made possible without the tremendous mentorship of Dr. Eng Tan (Emeritus: The Scripps Research Institute: see: Fritzler MJ, Chan EKL. Dr Eng M. Tan: a tribute to an enduring legacy in autoimmunity. Lupus 2017;26:208–217.). In addition, the author has had outstanding collaborators from around the world and the benefits of extremely brilliant graduate students and postdoctoral fellows. The author apologizes to colleagues whose works are not cited in this short perspective owing to publication limits.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.054 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.002 | 0.016 |
| Scholarly communication | 0.010 | 0.018 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.012 | 0.031 |
| Insufficient payload (model declined to judge) | 0.009 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".