Digital Contact Tracing Apps for COVID-19: Development of a Citizen-Centered Evaluation Framework
Bibliographic record
Abstract
Background The silent transmission of COVID-19 has led to an exponential growth of fatal infections. With over 4 million deaths worldwide, the need to control and stem transmission has never been more critical. New COVID-19 vaccines offer hope. However, administration timelines, long-term protection, and effectiveness against potential variants are still unknown. In this context, contact tracing and digital contact tracing apps (CTAs) continue to offer a mechanism to help contain transmission, keep people safe, and help kickstart economies. However, CTAs must address a wide range of often conflicting concerns, which make their development/evolution complex. For example, the app must preserve citizens’ privacy while gleaning their close contacts and as much epidemiological information as possible. Objective In this study, we derived a compare-and-contrast evaluative framework for CTAs that integrates and expands upon existing works in this domain, with a particular focus on citizen adoption; we call this framework the Citizen-Focused Compare-and-Contrast Evaluation Framework (C3EF) for CTAs. Methods The framework was derived using an iterative approach. First, we reviewed the literature on CTAs and mobile health app evaluations, from which we derived a preliminary set of attributes and organizing pillars. These attributes and the probing questions that we formulated were iteratively validated, augmented, and refined by applying the provisional framework against a selection of CTAs. Each framework pillar was then subjected to internal cross-team scrutiny, where domain experts cross-checked sufficiency, relevancy, specificity, and nonredundancy of the attributes, and their organization in pillars. The consolidated framework was further validated on the selected CTAs to create a finalized version of C3EF for CTAs, which we offer in this paper. Results The final framework presents seven pillars exploring issues related to CTA design, adoption, and use: (General) Characteristics, Usability, Data Protection, Effectiveness, Transparency, Technical Performance, and Citizen Autonomy. The pillars encompass attributes, subattributes, and a set of illustrative questions (with associated example answers) to support app design, evaluation, and evolution. An online version of the framework has been made available to developers, health authorities, and others interested in assessing CTAs. Conclusions Our CTA framework provides a holistic compare-and-contrast tool that supports the work of decision-makers in the development and evolution of CTAs for citizens. This framework supports reflection on design decisions to better understand and optimize the design compromises in play when evolving current CTAs for increased public adoption. We intend this framework to serve as a foundation for other researchers to build on and extend as the technology matures and new CTAs become available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".