Potential Pitfalls of Remote and Automated Video Assessments of Movements Disorders
Notice bibliographique
Résumé
Videoconferencing platform-based virtual clinical visits by neurologists have seen remarkable growth. Even postpandemic, it is anticipated that many patients with physical and cognitive impairment, particularly from remote areas, are likely to prefer remote monitoring. However, remote assessment has its downsides: Parkinson's disease, in particular, requires careful visual inspection of movements to establish the phenomenology and severity of movements.1 The increased demand for remote monitoring has fortunately coincided with the development of advanced artificial intelligence algorithms, suggesting that real-time automatic assessment of some disease aspects (eg, bradykinesia, tremor, hypomimia, altered eye movements) is a possibility. The potential to automate and more accurately quantify aspects of the semisubjective and time-consuming Unified Parkinson's Disease Rating Scale is tantalizing. However, the extreme sensitivity of "deep learning" automated algorithms to detect subtle changes in video recordings can be considered a two-edged sword: on the one hand, automated techniques may be able to detect clinical changes more sensitively than human experts; on the other hand, they also can be inappropriately sensitive to small perturbations that would be reasonably ignored by a human rater.2 Such sensitivity of deep learning models to "adversarial attacks" is well known in the machine learning community but may be less known in the clinical community, which may be increasingly dependent on automated diagnostic and disease-staging methods. Little is known about how real-life variations in frame rate, video resolution changes, and delays through the network can affect the ultimate automated rating relying on videoconferencing platforms. In a Zoom (https://zoom.us, Zoom Video Communications, Inc., San Jose, CA, USA) meeting of a group of 18 healthy adults physically located throughout the province of British Columbia, we simultaneously recorded finger-tapping movements under various resolution and network environments. The meeting latency statistics provided by Zoom were also documented by each participant. Four (22%) and 14 (78%) participants were connected via Ethernet and Wi-Fi, respectively, with baseline internet download speeds in the range of 124 to 918 Mbps. First, one subject recorded herself by a local camera application on a Mac (Apple Inc., Cupertino, CA, USA) laptop using the native "Photo Booth" app, at 1080p resolution and 60 frames/s to act as "ground truth." The subject simultaneously recorded herself through Zoom using the "pin" and "local recording" options of Zoom ("local Zoom recording"), resulting in a reduction of resolution and frame rates (720p × 25 frames/s). Other participants in the meeting also pinned the subject and recorded the Zoom display to their local devices as the "remote Zoom recording" (Zoom latency range, 61–84 ms). Lastly, a portion of the group recorded the gallery view of the meeting (5 × 4 grid) without pinning, and the subject was later cropped from the grid as the "remote cropped Zoom recording" (Zoom latency range, 61–352 ms), resulting in a substantial decline in resolution (Fig. 1A). We then compared the effects of these different recordings on a deep learning model of hand movements. In particular, the distance between the thumb and index finger was plotted from each video with a deep learning hand-pose estimation library (MediaPipe; Google, Mountain View, CA, USA). All analyses were performed with Python version 3.10.7 (Python Software Foundation, Wilmington, DE, USA). The local Zoom recording of the subject was reasonably accurate (Fig. 1B), but the accuracy of remote Zoom recordings varied substantially among the recorders, presumably because of individual network stability and hardware differences (Fig. 1C,D). The accuracy of finger movement tracking was least consistent in the remote, cropped Zoom recording due to the limited number of pixels to extract the finger landmarks, leading to significant errors (Fig. 1E). Our findings show that remote video can vary considerably depending on many realistic real-world factors. Although it is unlikely that these minor inconsistencies will significantly impact a human rater, care should be taken when relying on automated methods that have not been made robust to such perturbations. For example, the "jitters" observed in all Zoom settings could be misinterpreted by automated algorithms as true halts or hesitations during finger tapping. In addition, the remote-cropped Zoom recording underestimated the finger-tapping amplitude substantially. These technical factors could lead to misclassification of the Unified Parkinson's Disease Rating Scale bradykinesia score and subsequent suboptimal therapeutic decisions. Video-based telemedicine is here to stay in movement disorders. Variability in video acquisition should be carefully addressed in future studies of video-based digital biomarkers for Parkinson's disease, particularly if data are acquired remotely. This study was supported by a Canadian Institutes of Health Research/National Science and Engineering Research Council Collaborative Health Research Project (grant CPG-163986). We thank the members of the Pacific Parkinson Research Centre and the Department of Electrical and Computer Engineering of the University of British Columbia for participating in this study. 1. Research project: A. Conception, B. Organization, C. Execution;2. Data analysis: A. Design, B. Execution, C. Review and Critique;3. Manuscript: A. Writing of the first draft, B. Review and critique. K.W.P.: 1A, 1C, 2C, 3A, 3BH.J.W.: 1B, 1C, 2A, 2B, 3BT.Y.: 2A, 2B, 2C, 3BR.M.: 1B, 1C, 3BM.S.M.: 1B, 1C, 3BM.J.M.: 1A, 1B, 1C, 2C, 3B K.W.P. received a research grant from the Korean Neurological Association. H.J.W., T.Y., R.M., and M.S.M. have nothing to disclose. M.J.M. is supported by the John Nichol Chair in Parkinson's Research, Canadian Institutes of Health Research grant PJT-175305, and CHRP grant CPG-163986. The data that support the findings of this study are available from the corresponding author upon reasonable request.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,019 | 0,080 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,003 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,003 | 0,003 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».