Users' Perceptions and Trust in AI in Direct-to-Consumer mHealth: Qualitative Interview Study
Bibliographic record
Abstract
BACKGROUND: The increasing use of direct-to-consumer artificial intelligence (AI)-enabled mobile health (AI-mHealth) apps presents an opportunity for more effective health management and monitoring and expanded mobile health (mHealth) capabilities. However, AI's early developmental stage has prompted concerns related to trust, privacy, informed consent, and bias, among others. While some of these concerns have been explored in early stakeholder research related to AI-mHealth, the broader landscape of considerations that hold ethical significance to users remains underexplored. OBJECTIVE: Our aim was to document and explore the perspectives of individuals who reported previous experience using mHealth apps and their attitudes and ethically salient considerations regarding direct-to-consumer AI-mHealth apps. METHODS: As part of a larger study, we conducted semistructured interviews via Zoom with self-reported users of mHealth apps (N=21). Interviews consisted of a series of open-ended questions concerning participants' experiences, attitudes, and values relating to AI-mHealth apps and were conducted until topic saturation was reached. We collaboratively reviewed the interview transcripts and developed a codebook consisting of 37 codes describing recurring or otherwise noteworthy sentiments that inductively arose from the data. A single coder coded all transcripts, and the entire team contributed to conventional qualitative analysis. RESULTS: Our qualitative analysis yielded 3 major categories and 9 subcategories encompassing participants' perspectives. Participants described attitudes toward the impact of AI-mHealth on users' health and personal data (ie, influences on health awareness and management, value for mental vs physical health use cases, and the inevitability of data sharing), influences on their trust in AI-mHealth (ie, endorsements and guidance from health professionals or health or regulatory organizations, attitudes toward technology companies, and reasonable but not necessarily explainable output), and their preferences relating to the amount and type of information that is shared by AI-mHealth apps (ie, the types of data that are collected, future uses of user data, and the accessibility of information). CONCLUSIONS: This paper provides additional context relating to a number of concerns previously posited or identified in the AI-mHealth literature, including trust, explainability, and information sharing, and revealed additional considerations that have not been previously documented, that is, users' differentiation between the value of AI-mHealth for physical and mental health use cases and their willingness to extend empathy to nonexplainable AI. To the best of our knowledge, this study is the first to apply an open-ended, qualitative descriptive approach to explore the perspectives of end users of direct-to-consumer AI-mHealth apps.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.034 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.007 | 0.009 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".