Integrating Feedback From Application Reviews Into Software Development
Bibliographic record
Abstract
In application (app) development, effectively harnessing user feedback is crucial for enhancing app quality and user feedback. However, the vast and unstructured nature of user reviews often complicates these efforts, posing challenges in accurately capturing and integrating this feedback into the development processes. We automate the classification of issues in app reviews and examine how these issues correlate with code quality metrics (code smells and bug reports) and development activities (additions, deletions, and time to merge in pull requests). We aim to provide evidence-based guidance for effectively prioritizing and addressing user feedback. Employing a Mining Software Repositories (MSR) approach, we gathered and analyzed reviews from seven open-source Android apps. We evaluated the efficacy of three machine learning models-Support Vector Machines (SVM), BERT, and a fine-tuned GPT-3.5-for classifying issues in app reviews. The GPT-3.5 model achieved the highest accuracy at 95.0%. We found statistically significant correlations between the classified issues, code quality metrics, and development activities. However, these relationships varied across applications, highlighting the complex relationship between user feedback and the development process. Our study highlights the effectiveness of automated tools in identifying and classifying feedback within app reviews. Our automated approach enhances developers' ability to manage feedback effectively and supports optimal resource allocation to improve app quality and user feedback.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.180 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.008 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".