On the Impact of Development Frameworks on Mobile Apps
Bibliographic record
Abstract
Cross-platform mobile app development frame-works allow developers to use a single codebase to develop apps targeting different platforms. As these frame-works provide distinct features and may impact the apps' quality, their selection must be done with care. Although many works evaluated mobile frame-works, there is no synthesis on these studies. In this paper, we present a Systematic Literature Review (SLR) on approaches that evaluated cross- platform frame-works. Our SLR covers 75 papers and provides insights on 1) the most studied frame-works, 2) the criteria used for evaluation, 3) the evaluation methods used and 4) the results of these evaluations. The SLR shows that prior works generally used a prototype app to evaluate the frame-works but none explored the impact of the frame-works on the app's code quality. Thus, we carried out a preliminary empirical study on 3,566 mobile apps to evaluate the impact of mobile frame-works on the number of bugs and code smells in apps. The results of the study on native Android and React Native indicate that the latter has fewer code smells than native Android apps. Native Android apps generally had worse quality considering the number of bugs and code smells.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.068 | 0.301 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.018 | 0.010 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.006 | 0.007 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".