An automated mobile app labeling framework based on primary motivations for smartphone use
Bibliographic record
Abstract
Purpose This paper aims to propose an automated mobile app labeling framework based on a novel app classification scheme that is aligned with users’ primary motivations for using smartphones. The study addresses the gaps in incorporating the needs of users and other context information in app classification as well as recommendation systems. Design/methodology/approach Based on a corpus of mobile app descriptions collected from Google Play store, this study applies extensive text analytics and topic modeling procedures to profile mobile apps within the categories of the classification scheme. Sufficient number of representative and labeled app descriptions are then used to train a classifier using machine learning algorithms, such as rule-based, decision tree and artificial neural network. Findings Experimental results of the classifiers show high accuracy in automatically labeling new apps based on their descriptions. The accuracy of the classification results suggests a feasible direction in facilitating app searching and retrieval in different Web-based usage environments. Research limitations/implications As a common challenge in textual data projects, the problem of data size and data quality issues exists throughout the multiple phases of experiments. Future research will extend the data collection scope in many aspects to address the issues that constrained the current experiments. Practical implications These empirical experiments demonstrate the feasibility of textual data analysis in profiling apps and user context information. This study also benefits app developers by improving app descriptions through a better understanding of user needs and context information. Finally, the classification framework can also guide practitioners in customizing products and services beyond mobile apps where context information and user needs play an important role. Social implications Given the widespread usage and applications of smartphones today, the proposed app classification framework will have broader implications to different Web-based application environments. Originality/value While there have been other classification approaches in the literature, to the best of the authors’ knowledge, this framework is the first study on building an automated app labeling framework based on primary motivations of smartphone usage.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".