Co-Designing a Smoking Cessation Chatbot: Focus Group Study of End Users and Smoking Cessation Professionals
Bibliographic record
Abstract
BACKGROUND: Our prototype smoking cessation chatbot, Quin, provides evidence-based, personalized support delivered via a smartphone app to help people quit smoking. We developed Quin using a multiphase program of co-design research, part of which included focus group evaluation of Quin among stakeholders prior to clinical testing. OBJECTIVE: This study aimed to gather and compare feedback on the user experience of the Quin prototype from end users and smoking cessation professionals (SCPs) via a beta testing process to inform ongoing chatbot iterations and refinements. METHODS: Following active and passive recruitment, we conducted web-based focus groups with SCPs and end users from Queensland, Australia. Participants tested the app for 1-2 weeks prior to focus group discussion and could also log conversation feedback within the app. Focus groups of SCPs were completed first to review the breadth and accuracy of information, and feedback was prioritized and implemented as major updates using Agile processes prior to end user focus groups. We categorized logged in-app feedback using content analysis and thematically analyzed focus group transcripts. RESULTS: In total, 6 focus groups were completed between August 2022 and June 2023; 3 for SCPs (n=9 participants) and 3 for end users (n=7 participants). Four SCPs had previously smoked, and most end users currently smoked cigarettes (n=5), and 2 had quit smoking. The mean duration of focus groups was 58 (SD 10.9; range 46-74) minutes. We identified four major themes from focus group feedback: (1) conversation design, (2) functionality, (3) relationality and anthropomorphism, and (4) role as a smoking cessation support tool. In response to SCPs' feedback, we made two major updates to Quin between cohorts: (1) improvements to conversation flow and (2) addition of the "Moments of Crisis" conversation tree. Participant feedback also informed 17 recommendations for future smoking cessation chatbot developments. CONCLUSIONS: Feedback from end users and SCPs highlighted the importance of chatbot functionality, as this underpinned Quin's conversation design and relationality. The ready accessibility of accurate cessation information and impartial support that Quin provided was recognized as a key benefit for end users, the latter of which contributed to a feeling of accountability to the chatbot. Findings will inform the ongoing development of a mature prototype for clinical testing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.031 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.008 | 0.004 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".