The Online Patient Satisfaction Index for Patients With Low Back Pain: Development, Reliability, and Validation Study
Bibliographic record
Abstract
BACKGROUND: Low back pain is highly prevalent, and most often, a specific causative factor cannot be identified. Therefore, for most patients, their low back pain is labeled as nonspecific. Patient education and information are recommended for all these patients. The internet is an accessible source of medical information on low back pain. Approximately 50% of patients with low back pain search the internet for health and medical advice. Patient satisfaction with education and information is important in relation to patients' levels of inclination to use web-based information and their trust in the information they find. Although patients who are satisfied with the information they retrieve use the internet as a supplementary source of information, dissatisfied patients tend to avoid using the internet. Consumers' loyalty to a product is often applied to evaluate their satisfaction. Consumers have been shown to be good ambassadors for a service when they are willing to recommend the service to a friend or colleague. When consumers are willing to recommend a service to a friend or colleague, they are also likely to be future users of the service. To the best of our knowledge, no multi-item instrument exists to specifically evaluate satisfaction with information delivered on the web for people with low back pain. OBJECTIVE: This study aims to report on the development, reliability testing, and construct validity testing of the Online Patient Satisfaction Index to measure patients' satisfaction with web-based information for low back pain. METHODS: This is a cross-sectional validation study of the Online Patient Satisfaction Index. The index was developed with experts and assessed for face validity. It was subsequently administered to 150 adults with nonspecific low back pain. Of these, 46% (70/150) were randomly assigned to participate in a reliability test using an intraclass correlation coefficient of agreement. Construct validity was evaluated by hypothesis testing based on a web app (MyBack) and Wikipedia on low back pain. RESULTS: The index includes 8 items. The median score (range 0-24) based on the MyBack website was 20 (IQR 18-22), and the median score for Wikipedia was 12 (IQR 8-15). The entire score range was used. Overall, 53 participants completed a retest, of which 39 (74%) were stable in their satisfaction with the home page and were included in the analysis for reliability. Intraclass correlation coefficient of agreement was estimated to be 0.82 (95% CI 0.68-0.90). Two hypothesized correlations for construct validity were confirmed through an analysis using complete data. CONCLUSIONS: The index had good face validity, excellent reliability, and good construct validity and can be used to measure satisfaction with the provision of web-based information regarding nonspecific low back pain among people willing to access the internet to obtain health information. TRIAL REGISTRATION: ClinicalTrials.gov NCT03449004; https://clinicaltrials.gov/ct2/show/NCT03449004.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".