Comparison of health utility values from EQ-5D-3L and EQ-5D-5L in patients with breast cancer in different health states.
Bibliographic record
Abstract
197 Background: Recently, our group reported Canadian-derived EQ-5D-3L-derived health utility scores for over 20 different cancer sites (PMID: 27567613). However, more recently, a Canadian valuation set for EQ-5D-5L derived health utilities has been reported. As we switch from the 3L to 5L version, we need to evaluate whether the data can be analyzed together across versions, or whether the older cancer valuations need to be replicated. Methods: Breast cancer patients in three health states (primary breast cancer, locoregional recurrence of breast cancer, and metastatic breast cancer) were evaluated using EQ-5D-3L from 2014-2015 and EQ-5D-5L from 2016-2017. Mean (SD) values were compared. We opted not to compare the two versions in the same patients because of recall bias, so we compared two different sets of patients. Results: Of 387 breast cancer patients, 259 (67%) had completed the EQ-5D-5L, while 128 (33%) completed the EQ-5D-3L. The two groups had similar distributions for clinico-demographic data except for ethnicity, where there were more Asian patients in the 5L (22% vs 9%; p < 0.0074). When comparing EQ-5D-5L and EQ-5L-3L within health states, the EQ-5D-5L values were numerically higher for all three health states: for primary breast cancer (mean HUS(SD)/version: 0.85(0.11)/5L vs. 0.81(0.16)/3L) with a p-value of 0.028; locoregional recurrence (0.87(0.08)/5L vs. 0.85(0.20)/5L) with no available p-value due to the small sample size; metastatic disease (0.79(0.15)/5L vs. 0.78(0.13)/3L) with a p-value of 0.78. Histogram distributions appeared different between the two versions. There was greater separation and statistical significance between health states when using the 5L version when compared to the 3L version. The 5L version had a greater variation of values than the 3L, utilizing the entire spectrum of values from 0.0 to 1.0. Conclusions: Though the trends across health states were similar in EQ-5D-5L and EQ-5D-3L health states, values using the 5L were higher and distributions appeared different. There should be extreme caution if considering combining health utility values from 5L and 3L in Canadian breast cancer datasets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.024 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".