Are Digital Health Tools Gender-Inclusive? A Case Study of Smoking Cessation Chatbots
Bibliographic record
Abstract
Context Digital health technologies such as chatbots have the potential to improve health outcomes. However, neglecting sex and gender considerations in their development may perpetuate or even exacerbate health inequities. Objective This study assessed whether sex and gender were considered in the development and testing of smoking cessation chatbots. It also explored expert perspectives to generate recommendations for more gender-inclusive digital health design. Study Design and Analysis We conducted a case study of 11 smoking cessation chatbot trials identified through a prior systematic review. Each study was evaluated for inclusion of women as participants, application of SGBA for study outcomes and user experience, use of co-design methods, and gender parity among authors. A multidisciplinary panel of 12 experts across tobacco addiction, sex/gender science, and digital health was surveyed and convened to develop recommendations. Setting The case study was conducted at the INTREPID Lab, Centre for Addiction and Mental Health (CAMH), Toronto, between October 2023 and May 2024. Population Studied Published chatbot trials targeting adults who smoke. Intervention No direct intervention was tested. Instead, the project evaluated existing chatbot studies and gathered qualitative input through expert consultation. Outcome Measures We assessed whether studies reported on sample characteristics, sex/gender-disaggregated results for effectiveness and user experience, co-design involvement, and author gender composition. Expert recommendations were synthesized based on survey and meeting data. Results Only 45% of chatbot studies reported sex-disaggregated effectiveness outcomes and 18% disaggregated user experience data. Co-design methods were used in 27% of studies. Women represented 36% of first authors and 9% of last authors. The panel emphasized the need for SGBA, interdisciplinary development teams, and tool co-design with users. Mandatory reporting of SGBA by journals and funders may improve gender-inclusivity of emerging digital tools. Conclusions Despite the promise of digital health tools, most smoking cessation chatbot studies lacked a SGBA and gender-inclusive design. This case study underscores the importance of including sex and gender considerations across all stages of digital tool development to enhance equity and clinical impact.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.051 | 0.084 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.014 | 0.009 |
| Scholarly communication | 0.008 | 0.010 |
| Open science | 0.003 | 0.008 |
| Research integrity | 0.005 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".