MétaCan
Menu
Retour à la cohorte
Enregistrement W4406275130 · doi:10.1111/add.16764

Commentary on Sun and Tang: Measurement assessment and validity in problematic smartphone use

2025· article· en· W4406275130 sur OpenAlexfundno aff
Richard J. E. James, Lucy Hitcham

Notice bibliographique

RevueAddiction · 2025
Typearticle
Langueen
DomaineSocial Sciences
ThématiqueImpact of Technology on Adolescents
Établissements canadiensnon disponible
Organismes subventionnairesEngineering and Physical Sciences Research CouncilGambling Research Exchange Ontario
Mots-clésPsychologyClinical psychology

Résumé

récupéré en direct d'OpenAlex

Digital technologies have been targeted for restrictions, partly based on their purported addictive tendencies. However, the field is plagued by measurement problems. This commentary explores some of the key lessons drawn from Sun and Tang's study, and the implications for measuring both problematic smartphone use and other addictive behaviours. The thoughtful choice of estimation procedures for the confirmatory factor analysis (CFA) and invariance testing is worth particular attention. Many assessment studies use maximum likelihood (ML or MLR with robust standard errors) for CFA despite well-known limitations when applied to ordinal data [7]. A popular alternative is to use limited information estimation, for example, weighted least squares (WLSMV) to overcome these. However, doing so comes with major drawbacks, most notably when assessing measurement invariance [8, 9]. Sun and Tang [5] carefully balance the strengths of both MLR and WLSMV to validate the Problematic Smartphone Use Scale among Chinese college students (PSUS-C). These considerations are valuable across the entirety of addiction research, especially in domains or populations where endorsement of indicators might be skewed (e.g. gambling, certain forms of substance use and general population samples). To illustrate why these problems matter, CFA studies have repeatedly shown inconsistent evidence of structural validity in prominent scales such as the Problem Gambling Severity Index [10, 11]. However, closer examination suggests that most of this inconsistency is an artifact of using ML on ordinal questionnaire items in general population samples where the distribution of responses is often skewed. When analyzed using an approach that balances the strengths of both ML and WLSMV, these inconsistencies disappear [11, 12]. The findings also highlight an important tension between identifying the best-fitting factor structure and deciding how a scale should be used. Both exploratory factor analysis (EFA) and CFA rejected a single-factor model in this study, yet a sum score was used to assess criterion validity. We raise this to promote the benefits of testing models specifying either a second-order or a bifactor structure because these can assess whether a single score is appropriate [13]. This is an issue across the PSU field, where many scales have been validated as multi-dimensional. but are used as a single score. This tension is a source of analytic flexibility and a potential threat to the validity of many findings, especially when methods such as structural equation modelling are used. Our final reflection underscores the importance of invariance testing. Despite concluding in favor of strict invariance, there does not appear to be a comparison of latent mean differences that would allow a stronger test of group differences. Our examination of the descriptive data suggests the absence of a substantial sex difference in PSUS-C scores in this large, externally representative sample. We calculated the standardized effect size (d) using the mean (M) and SD statistics reported in table 1 (men: M = 58.05, SD = 18.09; women: M = 57.52, SD = 16.12). The difference observed in this study does not appear to practically differ from zero (d = 0.03). This finding contrasts with a large, disparate literature that has inconsistently found sex differences in the severity and prevalence of problematic smartphone behaviors (e.g. Cohen's d for women > men = 0.16 [14], 0.39 [15], 0.22 [16], 0.10 [17] and 0.21 [17]). This is further complicated by a fixation on creating novel instruments or adapting scales from other behavioral addictions instead of improving and refining existing measures [18]. Ultimately, the absence of appropriate psychometric validation found in many PSU and behavioral addiction measures makes it impossible to determine whether the group differences observed elsewhere reflect genuine differences or bias caused by sampling, specific measurement scales or specific questionnaire items. Sun and Tang's study [5] offers insights on how to move forward with the assessment and validation of behavioral addiction measurement scales. The use of rigorous testing is essential to establish whether addiction constructs are equivalent across diverse groups of people to make valid group comparisons and inferences [19]. Richard J. E. James: Writing—original draft (equal). Lucy Hitcham: Writing—original draft (equal). None. R.J. has received funding for gambling research projects in the last 3 years from GREO Evidence Insights and the Academic Forum for the Study of Gambling. These funds are sourced from regulatory settlements levied by the Gambling Commission in lieu of penalties. L.H. is funded by the Engineering and Physical Sciences Research Council (EPSRC) on a PhD Studentship scholarship (EP/S023305/1). Data sharing is not applicable to this article as no new data were created or analyzed in this study.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,069
Score d'incertitude au seuil0,278

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,057
Tête enseignante GPT0,340
Écart entre enseignants0,283 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueAddictionMême sujetImpact of Technology on AdolescentsTravaux en français237 207