Psychological and Behavioral Insights From Social Media Users: Natural Language Processing–Based Quantitative Study on Mental Well-Being
Notice bibliographique
Résumé
Background Depression significantly impacts an individual’s thoughts, emotions, behaviors, and moods; this prevalent mental health condition affects millions globally. Traditional approaches to detecting and treating depression rely on questionnaires and personal interviews, which can be time consuming and potentially inefficient. As social media has permanently shifted the pattern of our daily communications, social media postings can offer new perspectives in understanding mental illness in individuals because they provide an unbiased exploration of their language use and behavioral patterns. Objective This study aimed to develop and evaluate a methodological language framework that integrates psychological patterns, contextual information, and social interactions using natural language processing and machine learning techniques. The goal was to enhance intelligent decision-making for detecting depression at the user level. Methods We extracted language patterns via natural language processing approaches that facilitate understanding contextual and psychological factors, such as affective patterns and personality traits linked with depression. Then, we extracted social interaction influence features. The resultant social interaction influence that users have within their online social group is derived based on users’ emotions, psychological states, and context of communication extracted from status updates and the social network structure. We empirically evaluated the effectiveness of our framework by applying machine learning models to detect depression, reporting accuracy, recall, precision, and F1-score using social media status updates from 1047 users along with their associated depression diagnosis questionnaire scores. These datasets also include user postings, network connections, and personality responses. Results The proposed framework demonstrates accurate and effective detection of depression, improving performance compared to traditional baselines with an average improvement of 6% in accuracy and 10% in F1-score. It also shows competitive performance relative to state-of-the-art models. The inclusion of social interaction features demonstrates strong performance. By using all influence features (affective influence features, contextual influence features, and personality influence features), the model achieved an accuracy of 77% and a precision of 80%. Using affective features and affective influence features also showed strong performance, achieving 81% precision and an F1-score of 79%. Conclusions The developed framework offers practical applications, such as accelerating hospital diagnoses, improving prediction accuracy, facilitating timely referrals, and providing actionable insights for early interventions in mental health treatment plans.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,014 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».