Mobile Health Apps for Breast Cancer: Content Analysis and Quality Assessment
Bibliographic record
Abstract
BACKGROUND: The number of mobile health apps is rapidly increasing. This means that consumers are faced with a bewildering array of choices, and finding the benefit of such apps may be challenging. The significant international burden of breast cancer (BC) and the potential of mobile health apps to improve medical and public health practices mean that such apps will likely be important because of their functionalities in daily life. As the app market has grown exponentially, several review studies have scrutinized cancer- or BC-related apps. However, those reviews concentrated on the availability of the apps and relied on user ratings to decide on app quality. To minimize subjectivity in quality assessment, quantitative methods to assess BC-related apps are required. OBJECTIVE: The purpose of this study is to analyze the content and quality of BC-related apps to provide useful information for end users and clinicians. METHODS: Based on a stepwise systematic approach, we analyzed apps related to BC, including those related to prevention, detection, treatment, and survivor support. We used the keywords "breast cancer" in English and Korean to identify commercially available apps in the Google Play and App Store. The apps were then independently evaluated by 2 investigators to determine their eligibility for inclusion. The content and quality of the apps were analyzed using objective frameworks and the Mobile App Rating Scale (MARS), respectively. RESULTS: The initial search identified 1148 apps, 69 (6%) of which were included. Most BC-related apps provided information, and some recorded patient-generated health data, provided psychological support, and assisted with medication management. The Kendall coefficient of concordance between the raters was 0.91 (P<.001). The mean MARS score (range: 1-5) of the apps was 3.31 (SD 0.67; range: 1.94-4.53). Among the 5 individual dimensions, functionality had the highest mean score (4.37, SD 0.42) followed by aesthetics (3.74, SD 1.14). Apps that only provided information on BC prevention or management of its risk factors had lower MARS scores than those that recorded medical data or patient-generated health data. Apps that were developed >2 years ago, or by individuals, had significantly lower MARS scores compared to other apps (P<.001). CONCLUSIONS: The quality of BC-related apps was generally acceptable according to the MARS, but the gaps between the highest- and lowest-rated apps were large. In addition, apps using personalized data were of higher quality than those merely giving related information, especially after treatment in the cancer care continuum. We also found that apps that had been updated within 1 year and developed by private companies had higher MARS scores. This may imply that there are criteria for end users and clinicians to help choose the right apps for better clinical outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.114 | 0.267 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.007 |
| Bibliometrics | 0.045 | 0.024 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".