Empirical Testing of the External Validity of a Discrete Choice Experiment to Determine Preferred Treatment Option: The Case of Sleep Apnea
Bibliographic record
Abstract
There is an increasing use of the discrete choice experiment (DCE) method in health care to estimate preferences of individuals and the public for different services. Despite this increasing use, there are few studies that investigate the validity of the DCE in health. This study investigates the external validity of DCE by comparing the predicted treatment choices from the DCE to the actual treatment choices made by the same respondents using a decision board (DB) approach. The sample includes 140 patients who came for a sleep apnea routine visit in a hospital setting. Each respondent answered 10 DCE tasks and 1 DB task. The preferences were estimated with a generalized multinomial logit model and the predicted and actual treatment choices were compared both at the sample and individual levels. The results raise questions about the external validity of DCE in health. At the sample level, the comparison showed large but not significant differences between the two methods. This can be explained in part by the aggregation process that obscures variability in the individuals' preferences. At the individual level, the comparison showed that the two methods led to significantly different patterns of choices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".