Using Wearable Devices and Speech Data for Personalized Machine Learning in Early Detection of Mental Disorders: Protocol for a Participatory Research Study
Notice bibliographique
Résumé
BACKGROUND: Early identification of mental disorder symptoms is crucial for timely treatment and reduction of recurring symptoms and disabilities. A tool to help individuals recognize warning signs is important. We posit that such a tool would have to rely on longitudinal analysis of patterns and trends in the individual's daily activities and mood, which can now be captured through data from wearable activity trackers, speech recordings from mobile devices, and the individual's own description of their mental state. In this paper, we describe such a tool developed by our team to detect early signs of depression, anxiety, and stress. OBJECTIVE: This study aims to examine three questions about the effectiveness of machine learning models constructed based on multimodal data from wearables, speech, and self-reports: (1) How does speech about issues of personal context differ from speech while reading a neutral text, what type of speech data are more helpful in detecting mental health indicators, and how is the quality of the machine learning models influenced by multilanguage data? (2) Does accuracy improve with longitudinal data collection and how, and what are the most important features? and (3) How do personalized machine learning models compare against population-level models? METHODS: We collect longitudinal data to aid machine learning in accurately identifying patterns of mental disorder symptoms. We developed an app that collects voice, physiological, and activity data. Physiological and activity data are provided by a variety of off-the-shelf fitness trackers, that record steps, active minutes, duration of sleeping stages (rapid eye movement, deep, and light sleep), calories consumed, distance walked, heart rate, and speed. We also collect voice recordings of users reading specific texts and answering open-ended questions chosen randomly from a set of questions without repetition. Finally, the app collects users' answers to the Depression, Anxiety, and Stress Scale. The collected data from wearable devices and voice recordings will be used to train machine learning models to predict the levels of anxiety, stress, and depression in participants. RESULTS: The study is ongoing, and data collection will be completed by November 2023. We expect to recruit at least 50 participants attending 2 major universities (in Canada and Mexico) fluent in English or Spanish. The study will include participants aged between 18 and 35 years, with no communication disorders, acute neurological diseases, or history of brain damage. Data collection complied with ethical and privacy requirements. CONCLUSIONS: The study aims to advance personalized machine learning for mental health; generate a data set to predict Depression, Anxiety, and Stress Scale results; and deploy a framework for early detection of depression, anxiety, and stress. Our long-term goal is to develop a noninvasive and objective method for collecting mental health data and promptly detecting mental disorder symptoms. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/48210.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,059 | 0,065 |
| Méta-épidémiologie (sens strict) | 0,004 | 0,004 |
| Méta-épidémiologie (sens large) | 0,005 | 0,005 |
| Bibliométrie | 0,003 | 0,003 |
| Études des sciences et des technologies | 0,009 | 0,004 |
| Communication savante | 0,004 | 0,004 |
| Science ouverte | 0,003 | 0,004 |
| Intégrité de la recherche | 0,006 | 0,010 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,050 | 0,015 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».