Integrating CEN ISO/TS 82304-2 in the Catalan Health App Assessment Framework: Comparative Case Study
Bibliographic record
Abstract
Background: Health apps are increasingly being used to promote health, manage diseases, and deliver health care services. Still, there is scarce objective information regarding their quality beyond the required Conformité Européenne mark for medical apps, leading to potential risks for users. To address these challenges, several authorities have developed health app assessment frameworks. In 2017, the TIC Salut Social Foundation (FTSS) in Catalonia developed its own health app assessment framework, which has been in use since that year. The publication of CEN ISO/TS 82304-2 (abbreviated as 82304-2)-a Technical Specification for assessing health apps-and the cocreation of the Label2Enable 82304-2 handbook for certified assessment organizations provide a unique opportunity to harmonize app assessments across the European Union. Objective: This study aimed to perform a comparative analysis of the FTSS assessment framework with 82304-2 to explore the integration of 82304-2 in Catalonia. Our broader aim was to provide this methodology for health authorities elsewhere to consider integrating 82304-2 or other evaluation frameworks. Methods: For the comparative analysis, a mixed methods approach was used, combining a qualitative case study with a quantitative analysis of the 2 frameworks. The qualitative evaluation covered rationale for assessment, framework characteristics, governance, workflows, quality aspects, and quality requirements. For the quantitative analysis, all FTSS and 82304-2 requirements were translated into concepts and subconcepts. A scoring system identified matches of the frameworks with these subconcepts, with scores ranging from 0 (no match) to 0.5 (partial match) and 1 (full match). Integration was evaluated considering several scenarios, including adopting the Label2Enable 82304-2 handbook, adopting the 82304-2 requirements, adapting the 82304-2 requirements to local needs, and maintaining the current FTSS framework. Results: The main difference between the frameworks was the app usage-based assessment (FTSS) versus evidence- and app usage-based assessment (82304-2). All 120 FTSS requirements and 74 quality aspect-related 82304-2 requirements were translated into 78 concepts and 97 subconcepts. Overall, 48% (47/97) of the subconcepts were found in both frameworks, 39% (37.5/97) were specific to 82304-2, and 13% (12.5/97) were specific to FTSS. All 82304-2-specific subconcepts and thus all 82304-2 requirements were found to be relevant to FTSS. FTSS decided to integrate (adopt and adapt) all 74 82304-2 requirements. In total, 5 FTSS-specific requirements were included in the Label2Enable 82304-2 handbook, while another 4 rigor-enhancing requirements, 1 scope-expanding requirement, and 1 context-specific requirement would be assessed on top. Conclusions: The comprehensive comparative analysis of the FTSS framework and 82304-2 enabled FTSS decision-making to integrate all 82304-2 quality requirements and adopt the Label2Enable 82304-2 handbook in the future. The many new and all relevant 82304-2 concepts, the rigor of the handbook, and the few remaining FTSS-specific requirements are expected to be indicative of 82304-2's potential to make harmonized, robust health app assessments common in Catalonia and elsewhere. FTSS encourages other authorities to perform a similar evaluation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.054 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.008 | 0.003 |
| Open science | 0.003 | 0.007 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".