Validity, reliability and responsiveness of digital visual analogue scales for chronic spontaneous urticaria monitoring: A <scp>CRUSE</scp> ® mobile health study
Bibliographic record
Abstract
BACKGROUND: CRUSE® is an app that allows patients with chronic spontaneous urticaria (CSU) to monitor their daily disease activity through the use of visual analogue scales (VASs). We aimed to determine the concurrent validity, reliability, responsiveness and minimal important difference (MID) of CRUSE® VASs. METHODS: We evaluated the properties of three daily VASs: VAS for how much patients were affected by their CSU ('VAS urticaria'), VAS for the impact of urticaria on work/school productivity ('VAS productivity') and the VAS of EQ-5D. Concurrent validity was assessed by measuring the association between each VAS and the Urticaria Activity Score (UAS). Intra-rater reliability was determined based on the data of users providing multiple daily questionnaires within the same day. Test-retest reliability and responsiveness (ability to change), respectively, were tested in clinically stable and clinically unstable users. MIDs were determined using distribution-based methods. RESULTS: We included 5938 patients (67,380 days). Concurrent validity was high, with VAS urticaria being more strongly associated with the UAS score than the remaining VASs. Intra-rater reliability was also high, with intraclass correlation coefficients (ICC) being above 0.950 for all VASs. Moderate-high test-retest reliability and responsiveness were observed, with reliability ICC being highest for VAS EQ-5D and responsiveness being highest for VAS urticaria. The MID for VAS urticaria was 17 (out of 100) units, compared to 15 units for VAS productivity and 11 units for VAS EQ-5D. CONCLUSION: Daily VASs for CSU available in the CRUSE® app display high concurrent validity and intra-rater reliability and moderate-high test-retest reliability and responsiveness.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".