In-person versus virtual administration of the American College of Rheumatology gold standard cognitive battery in systemic lupus erythematosus: Are they interchangeable?
Bibliographic record
Abstract
OBJECTIVE: During the COVID-19 pandemic, many research studies were adapted, including our longitudinal study examining cognitive impairment (CI) in systemic lupus erythematosus (SLE). Cognitive testing was switched from in-person to virtual. This analysis aimed to determine if the administration method (in-person vs. virtual) of the ACR-neuropsychological battery (ACR-NB) affected participant cognitive performance and classification. METHODS: Data from our multi-visit, SLE CI study included demographic, clinical, and psychiatric characteristics, and the modified ACR-NB. Three analyses were undertaken for cognitive performance: (1) all visits, (2) non-CI group visits only and (3) intra-individual comparisons. A retrospective preferences questionnaire was given to participants who completed the ACR-NB both in-person and virtually. RESULTS: We analysed 328 SLE participants who had 801 visits (696 in-person and 105 virtual). Demographic, clinical, and psychiatric characteristics were comparable except for ethnicity, anxiety and disease-related damage. Across all three comparisons, six tests were consistently statistically significantly different. CI classification changed in 11/71 (15%) participants. 45% of participants preferred the virtual administration method and 33% preferred in-person. CONCLUSIONS: Of the 19 tests in the ACR-NB, we identified one or more problems with eight (42%) tests when moving from in-person to virtual administration. As the use of virtual cognitive testing will likely increase, these issues need to be addressed - potentially by validating a virtual version of the ACR-NB. Until then, caution must be taken when directly comparing virtual to in-person test results. If future studies use a mixed administration approach, this should be accounted for during analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".