Assessing the responsiveness of measures of oral health‐related quality of life
Bibliographic record
Abstract
OBJECTIVES: This paper illustrates ways of assessing the responsiveness of measures of oral health-related quality of life (OHRQoL) by examining the sensitivity of the oral health impact profile (OHIP)-14 to change when used to evaluate a dental care program for the elderly. METHODS: One hundred and sixteen elderly patients attending four municipally funded dental clinics completed a copy of the OHIP-14 prior to treatment and 1 month after the completion of treatment. The post-treatment questionnaire also included a global transition judgement that assessed subjects' perceptions of change in their oral health following treatment at the clinics. Change scores were calculated by subtracting post-treatment OHIP-14 scores from pre-treatment scores. The longitudinal construct validity of these change scores were assessed by means of their association with the global transition judgements. Measures of responsiveness included effect sizes for the change scores, the minimal important difference, and Guyatt's responsiveness index. An receiver operating characteristic (ROC) curve was constructed to determine the accuracy of the change scores in predicting whether patients had improved or not as a result of the treatment. RESULTS: Based on the global transition judgements, 60.2% of subjects reported improved oral health, 33.6% reported no change, and only 6.2% reported that it was a little worse. These changes are reflected in mean pre- and post-treatment OHIP-14 scores that declined from 15.8 to 11.5 (P < 0.001). Mean change scores showed a consistent gradient in the expected direction across categories of the global transition judgement, but differences between the groups were not significant. However, paired t-tests showed no significant differences in the pre- and post-treatment scores of stable subjects, but showed significant declines for subjects who reported improvement. Analysis of data from stable subjects indicated that OHIP-14 had excellent test-retest reliability with an intraclass correlation coefficient (ICC) of 0.84. Effect size based on change scores for all subjects and subgroups of subjects were small to moderate. The ROC analysis indicated that OHIP-14 change scores were not good "diagnostic tests" of improvement. The minimal important difference for the OHIP-14 was of 5-scale points, but detecting this difference would require relatively large sample sizes. CONCLUSIONS: OHIP-14 appeared to be responsive to change. However, the magnitude of change that it detected in the context described here was modest, probably because it was designed primarily as a discriminative measure. The psychometric properties of the global transition judgements that often provide the "gold standard" for responsiveness studies need to be established.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.106 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".