Epidemiology of discoid lupus erythematosus among adults in the United States: a cross‐sectional analysis
Bibliographic record
Abstract
Discoid lupus erythematosus (DLE) is a chronic cutaneous lupus erythematosus (CLE) characterized by well-demarcated, erythematous plaques, most commonly in the head and neck region.1 Prior estimates of DLE prevalence, ranging from 3.1 to 7.0 per 100,000, have been constrained by limited cohorts in urban areas with restricted patient demographic data.2, 3 We aim to ameliorate this gap via the National Institute of Health's All of Us research program, which prioritizes the inclusion of groups historically underrepresented in biomedical research. We specifically provide epidemiologic characteristics of DLE within the All of Us database, stratified by sociodemographic groups. We performed a cross-sectional analysis on the All of Us dataset on all patients with available electronic health record data. Using SNOMED code 200938002, we identified 922 DLE cases. Prevalence was estimated by calculating the proportion of DLE cases among all living patients with electronic health record (EHR) data. Univariable and multivariable logistic regression analyses were performed to determine the odds ratios (OR) of HS diagnosis within each patient category, with multivariable regression analysis controlled for demographic variables. 95% confidence intervals were calculated using the Wald method. Electronic health record data was retrieved for 287,011 patients, of which 922 had DLE. Table 1 reveals a greater than 5:1 ratio of female (0.46%, 95% CI 0.43–0.50) to male patients (0.09%, 95% CI 0.08–0.11). The prevalence of DLE was lowest among the oldest demographic, individuals aged 75 and above, with a prevalence of 0.22% (95% CI 0.18–0.27). Hispanic patients also had a higher prevalence of DLE at 0.37% (95% CI 0.32–0.42) compared to non-Hispanic patients. Among racial groups, African American patients exhibited the highest DLE prevalence (0.52%, 95% CI 0.46–0.58). Patients in the age range of 40–59 (aOR 1.75, 95% CI 1.45–2.11), individuals of Asian descent (aOR 1.65, 95% CI 1.11–2.37), and patients with an annual income between $10,000 and $25,000 (aOR 1.32, 95% CI 1.05–1.66) also exhibited heightened adjusted DLE odds in comparison to counterparts referenced in Table 2. Regarding education, those who attended college without earning a degree showed notably higher adjusted odds for DLE (aOR 1.75, 95% CI 1.35–2.28) compared to those with no high school degree (Table 2). Our study found a prevalence higher than previous estimates of 3.1 to 7.0 per 100,000 which were based on more restricted urban cohorts. This increase may reflect the broader diversity within the All of Us database. Furthermore, the higher prevalence of DLE among female patients is consistent with previous research, affirming the validity of our findings.2 However, current research is limited in explaining these phenomena within the DLE subclass. We, therefore, lean on the previously studied positive influences of cumulative smoke exposure and the X-linked genes, TLR7 and VGLL3, on the general CLE class as rationales behind the higher DLE affliction in older and female patients, respectively.4 Our results also suggest an association of DLE with reduced income levels. The impact of socioeconomic status has been researched concerning systemic lupus erythematosus and thus may extend to the DLE subclass.5 These findings underscore the presence of significant demographic disparities in DLE prevalence, emphasizing the importance of targeted screening and interventions for populations at elevated risk. Limitations include restricting the dataset to patients with EHRs and the potential for misclassification bias when using diagnostic codes to identify DLE cases. Furthermore, although diverse, this database is not a random sample of the U.S. population which may affect the generalizability of these findings. We attempt to limit such errors by combining a large, diverse sample size with controlling for confounding factors in statistical analysis but strongly recommend additional epidemiologic studies to further characterize DLE. The All of Us Research Program is supported by the National Institutes of Health, Office of the Director: Regional Medical Centers: 1 OT2 OD026549; 1 OT2 OD026554; 1 OT2 OD026557; 1 OT2 OD026556; 1 OT2 OD026550; 1 OT2 OD 026552; 1 OT2 OD026553; 1 OT2 OD026548; 1 OT2 OD026551; 1 OT2 OD026555; IAA #: AOD 16037; Federally Qualified Health Centers: HHSN 263201600085U; Data and Research Center: 5 U2C OD023196; Biobank: 1 U24 OD023121; The Participant Center: U24 OD023176; Participant Technology Systems Center: 1 U24 OD023163; Communications and Engagement: 3 OT2 OD023205; 3 OT2 OD023206; and Community Partners: 1 OT2 OD025277; 3 OT2 OD025315; 1 OT2 OD025337; 1 OT2 OD025276. In addition, the All of Us Research Program would not be possible without the partnership of its participants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.005 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".