Usability and Usefulness of a Symptom Management Coaching System for Patients With Cancer Treated With Immune Checkpoint Inhibitors: Comparative Mixed Methods Study
Bibliographic record
Abstract
BACKGROUND: The prognosis for patients with several types of cancer has substantially improved following the introduction of immune checkpoint inhibitors, a novel type of immunotherapy. However, patients may experience symptoms both from the cancer itself and from the medication. A prototype of the eHealth tool Cancer Patients Better Life Experience (CAPABLE) was developed to facilitate symptom management, aimed at patients with melanoma and renal cell carcinoma treated with immunotherapy. Better usability of such eHealth tools can lead to improved user well-being and reduced risk of harm. It is unknown for usability evaluations whether certain usability problems would only be evident to patients whose condition closely resembles the target population, or if a broader group of patients would lead to the identification of a broader range of potential usability issues. OBJECTIVE: This study aims to evaluate the CAPABLE prototype by conducting tests to assess usability, user experience, and perceived acceptability among end users, and to assess any agreements or differences in the results of our wide range of participants. METHODS: This usability study was executed by interviewing participants with a melanoma or renal cell carcinoma diagnosis who have received immunotherapy and participants without direct experience with the targeted cancer types who have not received immunotherapy. Participants were asked to review the concept of the tool, perform think-aloud tasks, and complete the System Usability Scale and a Perceived Usefulness questionnaire. Usability problems were extracted from the interview data by independent coding and mapped to an eHealth Usability Problem Framework. RESULTS: We included 21 participants in the study, aged 29 to 73 years; 13 participants who had received immunotherapy and 8 participants who had not received immunotherapy. In total, 76 usability problems were identified. A total of 22 usability problems were in the task-technology fit category of the usability framework, mostly regarding the coaching and symptom functionality of the prototype. Critical problems regarding the symptom monitoring functionality were mainly found by participants who had received immunotherapy. For 8 out of 10 statements in the Perceived Usefulness questionnaire, more than 75% of participants agreed or strongly agreed. The overall mean System Usability Scale score was 80 out of 100 (SD 11.3). CONCLUSIONS: Despite identified usability issues, participants responded positively to the Perceived Usefulness questionnaire regarding the evaluated tool. Further analysis of the usability problems indicates that it was essential to include participants who matched the target end users. Participants treated with immunotherapy, specifically with previous experience in immune-related adverse events, encountered critical problems with symptom reporting that would not have been identified if these participants were not included. For other tasks and functionalities, it seems likely that loosening the inclusion criteria would have resulted in sufficient feedback without critical missing usability issues.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.032 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".