MétaCan
Menu
Back to cohort
Record W4312019381 · doi:10.2196/43905

What Makes a Quality Health App—Developing a Global Research-Based Health App Quality Assessment Framework for CEN-ISO/TS 82304-2: Delphi Study

2022· article· en· W4312019381 on OpenAlexvenueno aff
Petra Hoogendoorn, Anke Versluis, Sanne van Kampen, Charles McCay, Matt Leahy, Marlou Bijlsma, Stefano Bonacina, Tobias Bonten, Marie-José Bonthuis, Anouk Butterlin, Koen Cobbaert, Thea Duijnhoven, Cynthia Hallensleben, Stuart Harrison, Mark Hastenteufel, Terhi Holappa, Ben Kokx, Birgit Morlion, Norbert Pauli, Frank Ploeg, Mark Salmon, Kyma Schnoor, Mary Sharp, Pier Angelo Sottile, Alpo Värri, Patricia Williams, Georg Heidenreich, Nicholas Oughtibridge, Robert Stegwee, Niels H. Chavannes

Bibliographic record

VenueJMIR Formative Research · 2022
Typearticle
Languageen
FieldHealth Professions
TopicMobile Health and mHealth Applications
Canadian institutionsnot available
FundersEuropean Commission
KeywordsStandardizationDelphi methodQuality (philosophy)CommissionDelphiProcess managementBusinessComputer sciencePolitical science

Abstract

fetched live from OpenAlex

BACKGROUND: The lack of an international standard for assessing and communicating health app quality and the lack of consensus about what makes a high-quality health app negatively affect the uptake of such apps. At the request of the European Commission, the international Standard Development Organizations (SDOs), European Committee for Standardization, International Organization for Standardization, and International Electrotechnical Commission have joined forces to develop a technical specification (TS) for assessing the quality and reliability of health and wellness apps. OBJECTIVE: This study aimed to create a useful, globally applicable, trustworthy, and usable framework to assess health app quality. METHODS: A 2-round Delphi technique with 83 experts from 6 continents (predominantly Europe) participating in one (n=42, 51%) or both (n=41, 49%) rounds was used to achieve consensus on a framework for assessing health app quality. Aims included identifying the maximum 100 requirement questions for the uptake of apps that do or do not qualify as medical devices. The draft assessment framework was built on 26 existing frameworks, the principles of stringent legislation, and input from 20 core experts. A follow-up survey with 28 respondents informed a scoring mechanism for the questions. After subsequent alignment with related standards, the quality assessment framework was tested and fine-tuned with manufacturers of 11 COVID-19 symptom apps. National mirror committees from the 52 countries that participated in the SDO technical committees were invited to comment on 4 working drafts and subsequently vote on the TS. RESULTS: The final quality assessment framework includes 81 questions, 67 (83%) of which impact the scores of 4 overarching quality aspects. After testing with people with low health literacy, these aspects were phrased as "Healthy and safe," "Easy to use," "Secure data," and "Robust build." The scoring mechanism enables communication of the quality assessment results in a health app quality score and label, alongside a detailed report. Unstructured interviews with stakeholders revealed that evidence and third-party assessment are needed for health app uptake. The manufacturers considered the time needed to complete the assessment and gather evidence (2-4 days) acceptable. Publication of CEN-ISO/TS 82304-2:2021 Health software - Part 2: Health and wellness apps - Quality and reliability was approved in May 2021 in a nearly unanimous vote by 34 national SDOs, including 6 of the 10 most populous countries worldwide. CONCLUSIONS: A useful and usable international standard for health app quality assessment was developed. Its quality, approval rate, and early use provide proof of its potential to become the trusted, commonly used global framework. The framework will help manufacturers enhance and efficiently demonstrate the quality of health apps, consumers, and health care professionals to make informed decisions on health apps. It will also help insurers to make reimbursement decisions on health apps.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.092
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Science and technology studies, Research integrity
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.599
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0920.002
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.000
Bibliometrics0.0010.006
Science and technology studies0.0250.001
Scholarly communication0.0000.001
Open science0.0020.002
Research integrity0.0000.008
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.515
GPT teacher head0.703
Teacher spread0.188 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations28
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Formative ResearchSame topicMobile Health and mHealth ApplicationsFrench-language works237,207