MétaCan
Menu
Back to cohort
Record W4406275130 · doi:10.1111/add.16764

Commentary on Sun and Tang: Measurement assessment and validity in problematic smartphone use

2025· article· en· W4406275130 on OpenAlexfundno aff
Richard J. E. James, Lucy Hitcham

Bibliographic record

VenueAddiction · 2025
Typearticle
Languageen
FieldSocial Sciences
TopicImpact of Technology on Adolescents
Canadian institutionsnot available
FundersEngineering and Physical Sciences Research CouncilGambling Research Exchange Ontario
KeywordsPsychologyClinical psychology

Abstract

fetched live from OpenAlex

Digital technologies have been targeted for restrictions, partly based on their purported addictive tendencies. However, the field is plagued by measurement problems. This commentary explores some of the key lessons drawn from Sun and Tang's study, and the implications for measuring both problematic smartphone use and other addictive behaviours. The thoughtful choice of estimation procedures for the confirmatory factor analysis (CFA) and invariance testing is worth particular attention. Many assessment studies use maximum likelihood (ML or MLR with robust standard errors) for CFA despite well-known limitations when applied to ordinal data [7]. A popular alternative is to use limited information estimation, for example, weighted least squares (WLSMV) to overcome these. However, doing so comes with major drawbacks, most notably when assessing measurement invariance [8, 9]. Sun and Tang [5] carefully balance the strengths of both MLR and WLSMV to validate the Problematic Smartphone Use Scale among Chinese college students (PSUS-C). These considerations are valuable across the entirety of addiction research, especially in domains or populations where endorsement of indicators might be skewed (e.g. gambling, certain forms of substance use and general population samples). To illustrate why these problems matter, CFA studies have repeatedly shown inconsistent evidence of structural validity in prominent scales such as the Problem Gambling Severity Index [10, 11]. However, closer examination suggests that most of this inconsistency is an artifact of using ML on ordinal questionnaire items in general population samples where the distribution of responses is often skewed. When analyzed using an approach that balances the strengths of both ML and WLSMV, these inconsistencies disappear [11, 12]. The findings also highlight an important tension between identifying the best-fitting factor structure and deciding how a scale should be used. Both exploratory factor analysis (EFA) and CFA rejected a single-factor model in this study, yet a sum score was used to assess criterion validity. We raise this to promote the benefits of testing models specifying either a second-order or a bifactor structure because these can assess whether a single score is appropriate [13]. This is an issue across the PSU field, where many scales have been validated as multi-dimensional. but are used as a single score. This tension is a source of analytic flexibility and a potential threat to the validity of many findings, especially when methods such as structural equation modelling are used. Our final reflection underscores the importance of invariance testing. Despite concluding in favor of strict invariance, there does not appear to be a comparison of latent mean differences that would allow a stronger test of group differences. Our examination of the descriptive data suggests the absence of a substantial sex difference in PSUS-C scores in this large, externally representative sample. We calculated the standardized effect size (d) using the mean (M) and SD statistics reported in table 1 (men: M = 58.05, SD = 18.09; women: M = 57.52, SD = 16.12). The difference observed in this study does not appear to practically differ from zero (d = 0.03). This finding contrasts with a large, disparate literature that has inconsistently found sex differences in the severity and prevalence of problematic smartphone behaviors (e.g. Cohen's d for women > men = 0.16 [14], 0.39 [15], 0.22 [16], 0.10 [17] and 0.21 [17]). This is further complicated by a fixation on creating novel instruments or adapting scales from other behavioral addictions instead of improving and refining existing measures [18]. Ultimately, the absence of appropriate psychometric validation found in many PSU and behavioral addiction measures makes it impossible to determine whether the group differences observed elsewhere reflect genuine differences or bias caused by sampling, specific measurement scales or specific questionnaire items. Sun and Tang's study [5] offers insights on how to move forward with the assessment and validation of behavioral addiction measurement scales. The use of rigorous testing is essential to establish whether addiction constructs are equivalent across diverse groups of people to make valid group comparisons and inferences [19]. Richard J. E. James: Writing—original draft (equal). Lucy Hitcham: Writing—original draft (equal). None. R.J. has received funding for gambling research projects in the last 3 years from GREO Evidence Insights and the Academic Forum for the Study of Gambling. These funds are sourced from regulatory settlements levied by the Gambling Commission in lieu of penalties. L.H. is funded by the Engineering and Physical Sciences Research Council (EPSRC) on a PhD Studentship scholarship (EP/S023305/1). Data sharing is not applicable to this article as no new data were created or analyzed in this study.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.069
Threshold uncertainty score0.278

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.057
GPT teacher head0.340
Teacher spread0.283 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueAddictionSame topicImpact of Technology on AdolescentsFrench-language works237,207