MétaCan
Menu
Back to cohort
Record W4318302282 · doi:10.1002/mds.29325

Potential Pitfalls of Remote and Automated Video Assessments of Movements Disorders

2023· letter· en· W4318302282 on OpenAlexafffundabout
Kye Won Park, Hsin Jui Wu, Tianze Yu, Ravneet Mahal, Maryam S. Mirian, Martin J. McKeown

Bibliographic record

VenueMovement Disorders · 2023
Typeletter
Languageen
FieldMedicine
TopicNeurological disorders and treatments
Canadian institutionsCanadian Sport Centre PacificUniversity of British Columbia
FundersCanadian Institutes of Health Research
KeywordsComputer scienceArtificial intelligenceMachine learningPhysical medicine and rehabilitationMedicine

Abstract

fetched live from OpenAlex

Videoconferencing platform-based virtual clinical visits by neurologists have seen remarkable growth. Even postpandemic, it is anticipated that many patients with physical and cognitive impairment, particularly from remote areas, are likely to prefer remote monitoring. However, remote assessment has its downsides: Parkinson's disease, in particular, requires careful visual inspection of movements to establish the phenomenology and severity of movements.1 The increased demand for remote monitoring has fortunately coincided with the development of advanced artificial intelligence algorithms, suggesting that real-time automatic assessment of some disease aspects (eg, bradykinesia, tremor, hypomimia, altered eye movements) is a possibility. The potential to automate and more accurately quantify aspects of the semisubjective and time-consuming Unified Parkinson's Disease Rating Scale is tantalizing. However, the extreme sensitivity of "deep learning" automated algorithms to detect subtle changes in video recordings can be considered a two-edged sword: on the one hand, automated techniques may be able to detect clinical changes more sensitively than human experts; on the other hand, they also can be inappropriately sensitive to small perturbations that would be reasonably ignored by a human rater.2 Such sensitivity of deep learning models to "adversarial attacks" is well known in the machine learning community but may be less known in the clinical community, which may be increasingly dependent on automated diagnostic and disease-staging methods. Little is known about how real-life variations in frame rate, video resolution changes, and delays through the network can affect the ultimate automated rating relying on videoconferencing platforms. In a Zoom (https://zoom.us, Zoom Video Communications, Inc., San Jose, CA, USA) meeting of a group of 18 healthy adults physically located throughout the province of British Columbia, we simultaneously recorded finger-tapping movements under various resolution and network environments. The meeting latency statistics provided by Zoom were also documented by each participant. Four (22%) and 14 (78%) participants were connected via Ethernet and Wi-Fi, respectively, with baseline internet download speeds in the range of 124 to 918 Mbps. First, one subject recorded herself by a local camera application on a Mac (Apple Inc., Cupertino, CA, USA) laptop using the native "Photo Booth" app, at 1080p resolution and 60 frames/s to act as "ground truth." The subject simultaneously recorded herself through Zoom using the "pin" and "local recording" options of Zoom ("local Zoom recording"), resulting in a reduction of resolution and frame rates (720p × 25 frames/s). Other participants in the meeting also pinned the subject and recorded the Zoom display to their local devices as the "remote Zoom recording" (Zoom latency range, 61–84 ms). Lastly, a portion of the group recorded the gallery view of the meeting (5 × 4 grid) without pinning, and the subject was later cropped from the grid as the "remote cropped Zoom recording" (Zoom latency range, 61–352 ms), resulting in a substantial decline in resolution (Fig. 1A). We then compared the effects of these different recordings on a deep learning model of hand movements. In particular, the distance between the thumb and index finger was plotted from each video with a deep learning hand-pose estimation library (MediaPipe; Google, Mountain View, CA, USA). All analyses were performed with Python version 3.10.7 (Python Software Foundation, Wilmington, DE, USA). The local Zoom recording of the subject was reasonably accurate (Fig. 1B), but the accuracy of remote Zoom recordings varied substantially among the recorders, presumably because of individual network stability and hardware differences (Fig. 1C,D). The accuracy of finger movement tracking was least consistent in the remote, cropped Zoom recording due to the limited number of pixels to extract the finger landmarks, leading to significant errors (Fig. 1E). Our findings show that remote video can vary considerably depending on many realistic real-world factors. Although it is unlikely that these minor inconsistencies will significantly impact a human rater, care should be taken when relying on automated methods that have not been made robust to such perturbations. For example, the "jitters" observed in all Zoom settings could be misinterpreted by automated algorithms as true halts or hesitations during finger tapping. In addition, the remote-cropped Zoom recording underestimated the finger-tapping amplitude substantially. These technical factors could lead to misclassification of the Unified Parkinson's Disease Rating Scale bradykinesia score and subsequent suboptimal therapeutic decisions. Video-based telemedicine is here to stay in movement disorders. Variability in video acquisition should be carefully addressed in future studies of video-based digital biomarkers for Parkinson's disease, particularly if data are acquired remotely. This study was supported by a Canadian Institutes of Health Research/National Science and Engineering Research Council Collaborative Health Research Project (grant CPG-163986). We thank the members of the Pacific Parkinson Research Centre and the Department of Electrical and Computer Engineering of the University of British Columbia for participating in this study. 1. Research project: A. Conception, B. Organization, C. Execution;2. Data analysis: A. Design, B. Execution, C. Review and Critique;3. Manuscript: A. Writing of the first draft, B. Review and critique. K.W.P.: 1A, 1C, 2C, 3A, 3BH.J.W.: 1B, 1C, 2A, 2B, 3BT.Y.: 2A, 2B, 2C, 3BR.M.: 1B, 1C, 3BM.S.M.: 1B, 1C, 3BM.J.M.: 1A, 1B, 1C, 2C, 3B K.W.P. received a research grant from the Korean Neurological Association. H.J.W., T.Y., R.M., and M.S.M. have nothing to disclose. M.J.M. is supported by the John Nichol Chair in Parkinson's Research, Canadian Institutes of Health Research grant PJT-175305, and CHRP grant CPG-163986. The data that support the findings of this study are available from the corresponding author upon reasonable request.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.019
metaresearch head score (Gemma)0.080
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.019
Threshold uncertainty score0.100

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0190.080
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0010.002
Scholarly communication0.0020.002
Open science0.0030.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.014
GPT teacher head0.288
Teacher spread0.275 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes3
Has abstractyes

Explore more

Same venueMovement DisordersSame topicNeurological disorders and treatmentsFrench-language works237,207