Content validity of global measures for at-work productivity in patients with rheumatic diseases: an international qualitative study
Bibliographic record
Abstract
OBJECTIVES: To identify from a patient's perspective, difficulties and differences in the comprehension of five global presenteeism measures in patients with inflammatory arthritis and OA across seven countries. METHODS: Seventy patients with a diagnosis of inflammatory arthritis or OA in paid employment were recruited from seven countries across Europe and Canada. Patients were randomly allocated to be cognitively debriefed on 3/5 global measures [Work Productivity Scale - Rheumatoid Arthritis, Work Productivity and Activity Impairment Questionnaire (WPAI), Work Ability Index, Quality and Quantity questionnaire, and WHO Health and Work Performance Questionnaire (HPQ)], with the WPAI debriefed in all patients as a standard measure of comparison between countries and patients. NVivo was used to code the data into four themes: construct and anchor, time recall, reference frame, and attribution. RESULTS: Discrepancies were found in the interpretation of the word performance (HPQ) between countries, with Romania and Sweden relating performance to sports rather than work. Seventy percent of patients considered that a 7-day recall (WPAI) can accurately represent how their disease affects work productivity. The compared to normal reference (Quality and Quantity questionnaire) was reportedly too ambiguous, and the comparison with colleagues (HPQ), made many feel uncomfortable. Overall, 29% of patients said the WPAI was the most relevant to them, making it the most favoured measure. CONCLUSION: Overall, patients across countries agree that the construct of work productivity in the last 7 days can accurately reflect the impact of disease while at work. Some current constructs to assess at-work productivity are not interchangeable between languages.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".