What Makes a Quality Health App—Developing a Global Research-Based Health App Quality Assessment Framework for CEN-ISO/TS 82304-2: Delphi Study
Bibliographic record
Abstract
BACKGROUND: The lack of an international standard for assessing and communicating health app quality and the lack of consensus about what makes a high-quality health app negatively affect the uptake of such apps. At the request of the European Commission, the international Standard Development Organizations (SDOs), European Committee for Standardization, International Organization for Standardization, and International Electrotechnical Commission have joined forces to develop a technical specification (TS) for assessing the quality and reliability of health and wellness apps. OBJECTIVE: This study aimed to create a useful, globally applicable, trustworthy, and usable framework to assess health app quality. METHODS: A 2-round Delphi technique with 83 experts from 6 continents (predominantly Europe) participating in one (n=42, 51%) or both (n=41, 49%) rounds was used to achieve consensus on a framework for assessing health app quality. Aims included identifying the maximum 100 requirement questions for the uptake of apps that do or do not qualify as medical devices. The draft assessment framework was built on 26 existing frameworks, the principles of stringent legislation, and input from 20 core experts. A follow-up survey with 28 respondents informed a scoring mechanism for the questions. After subsequent alignment with related standards, the quality assessment framework was tested and fine-tuned with manufacturers of 11 COVID-19 symptom apps. National mirror committees from the 52 countries that participated in the SDO technical committees were invited to comment on 4 working drafts and subsequently vote on the TS. RESULTS: The final quality assessment framework includes 81 questions, 67 (83%) of which impact the scores of 4 overarching quality aspects. After testing with people with low health literacy, these aspects were phrased as "Healthy and safe," "Easy to use," "Secure data," and "Robust build." The scoring mechanism enables communication of the quality assessment results in a health app quality score and label, alongside a detailed report. Unstructured interviews with stakeholders revealed that evidence and third-party assessment are needed for health app uptake. The manufacturers considered the time needed to complete the assessment and gather evidence (2-4 days) acceptable. Publication of CEN-ISO/TS 82304-2:2021 Health software - Part 2: Health and wellness apps - Quality and reliability was approved in May 2021 in a nearly unanimous vote by 34 national SDOs, including 6 of the 10 most populous countries worldwide. CONCLUSIONS: A useful and usable international standard for health app quality assessment was developed. Its quality, approval rate, and early use provide proof of its potential to become the trusted, commonly used global framework. The framework will help manufacturers enhance and efficiently demonstrate the quality of health apps, consumers, and health care professionals to make informed decisions on health apps. It will also help insurers to make reimbursement decisions on health apps.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.092 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.006 |
| Science and technology studies | 0.025 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.008 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".