Recall of health-related quality of life: how does memory affect the SF-6D in patients with psoriasis or multiple sclerosis? A prospective observational study in Germany
Bibliographic record
Abstract
OBJECTIVE: This study aimed to quantify recall bias in the measurement of health-related quality of life (HRQoL), that is, the extent to which recollection is impaired and leads to distorted judgements. DESIGN: Prospective observational study. SETTING AND PARTICIPANTS: One hundred patients with two paradigmatic chronic diseases (50 with multiple sclerosis and 50 with psoriasis) were recruited at two outpatient clinics. METHODS AND OUTCOME MEASURES: Patients completed the online version of the 12-Item Short Form Survey (SF-12) repeatedly for 28 consecutive days: (1) daily, considering the past 24 hours; (2) weekly, considering the past 7 days; and (3) on the last day of data collection, considering the past 4 weeks. SF-12 scores for all three measurement approaches were subsequently converted into preference-based utility indices (Short-Form Six-Dimension). Agreement of the three indices was analysed on group and individual patient levels. RESULTS: The mean age of participants was 40.3 years (±12.0), and 63% were female. The utility index based on daily recall (0.74±0.13) was more positive than indices based on a weekly (0.70±0.13, p<0.001) or a monthly (0.70±0.14, p<0.001) recall. While agreement of measurement approaches was high on group level (intraclass correlation coefficient>0.85), it was lower for the subgroup of patients experiencing high variability of HRQoL over time. Bland-Altman plots revealed considerable differences on individual patient level. CONCLUSIONS: On the group level, retrospective overestimation and underestimation of HRQoL almost cancelled out one another and recall bias was relatively small. Therefore, a 4-week recall period could be appropriate when group-level data are used for research or economic evaluations. In contrast, recall bias can be considerable on the individual patient level and may thus impact decision-making in clinical practice. TRIAL REGISTRATION NUMBER: VfD_RECALL_16_003837.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".