The accuracy of pain and fatigue items across different reporting periods
Bibliographic record
Abstract
The length of the reporting period specified for items assessing pain and fatigue varies among instruments. How the length of recall impacts the accuracy of symptom reporting is largely unknown. This study investigated the accuracy of ratings for reporting periods ranging from 1 day to 28 days for several items from widely used pain and fatigue measures (SF36v2, Brief Pain Inventory, McGill Pain Questionnaire, Brief Fatigue Inventory). Patients from a community rheumatology practice (N=83) completed momentary pain and fatigue items on average of 5.4 times per day for a month using an electronic diary. Averaged momentary ratings formed the basis for comparison with recall ratings interspersed throughout the month referencing 1-day, 3-day, 7-day, and 28-day periods. As found in previous research, recall ratings were consistently inflated relative to averaged momentary ratings. Across most items, 1-day recall corresponded well to the averaged momentary assessments for the day. Several, but not all, items demonstrated substantial correlations across the different reporting periods. An additional 7 day-by-day recall task suggested that patients have increasing difficulty actually remembering symptom levels beyond the past several days. These data were collected while patients were receiving usual care and may not generalize to conditions where new interventions are being introduced and outcomes evaluated. Reporting periods can influence the accuracy of retrospective symptom reports and should be a consideration in study design.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.080 | 0.217 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".