Do visual analogue scale (VAS) derived standard gamble (SG) utilities agree with Health Utilities Index utilities? A comparison of patient and community preferences for health status in rheumatoid arthritis patients
Bibliographic record
Abstract
BACKGROUND: Assessment of Health Related Quality of Life (HRQL) has become increasingly important and various direct and indirect methods and instruments have been devised to measure it. In direct methods such as Visual Analog Scale (VAS) and Standard Gamble (SG), respondent both assesses and values health states therefore the final score reflects patient's preferences. In indirect methods such as multi-attribute health status classification systems, the patient provides the assessment of a health state and then a multi-attribute utility function is used for evaluation of the health state. Because these functions have been estimated using valuations of general population, the final score reflects community's preferences. The objective of this study is to assess the agreement between community preferences derived from the Health Utilities Index Mark 2 (HUI2) and Mark 3 (HUI3) systems, and patient preferences. METHODS: Visual analog scale (VAS) and HUI scores were obtained from a sample of 320 rheumatoid arthritis patients. VAS scores were adjusted for end-aversion bias and transformed to standard gamble (SG) utility scores using 8 different power conversion formulas reported in other studies. Individual level agreement between SG utilities and HUI2 and HUI3 utilities was assessed using the intraclass correlation coefficient (ICC). Group level agreement was assessed by comparing group means using the paired t-test. RESULTS: After examining all 8 different SG estimates, the ICC (95% confidence interval) between SG and HUI2 utilities ranged from 0.45 (0.36 to 0.54) to 0.55 (0.47 to 0.62). The ICC between SG and HUI3 utilities ranged from 0.45 (0.35 to 0.53) to 0.57 (0.49 to 0.64). The mean differences between SG and HUI2 utilities ranged from 0.10 (0.08 to 0.12) to 0.22 (0.20 to 0.24). The mean differences between SG and HUI3 utilities ranged from 0.18 (0.16 to 0.2) to 0.28 (0.26 to 0.3). CONCLUSION: At the individual level, patient and community preferences show moderate to strong agreement, but at the group level they have clinically important and statistically significant differences. Using different sources of preference might alter clinical and policy decisions that are based on methods that incorporate HRQL assessment. VAS-derived utility scores are not good substitutes for HUI scores.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.091 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".