Enhancing Reliability of Automated Remote Parkinson's Assessments: Real‐World Video Quality Challenges
Bibliographic record
Abstract
We would like to share our experience with remotely capturing video recordings for partial motor assessment of Parkinson's disease (PD) as part of a national registry initiative. Remote, automated assessment of motor symptoms is becoming increasingly important for detecting symptom fluctuations and tracking disease progression in Parkinson's disease. Video-based assessments provide a scalable and objective approach to motor testing, enhancing accessibility, particularly for individuals living far from specialized Movement Disorder clinics.1 However, their reliability depends heavily on video quality, which is often compromised when recordings are captured “in the wild” without supervision or technical assistance. We analyzed 287 videos from 57 participants recruited from the Canadian Open Parkinson Network (C-OPN) Registry,2 collected as part of a project developing algorithms for automated assessment of motor tasks. Participants recorded the videos at home using their personal webcams, typically integrated into laptops or desktop computers. Both the task design and scoring rubric were based on the MDS-UPDRS Part III3 and included finger tapping, hand movements, pronation-supination, and assessment of tremor.4 Despite unambiguous instructions (Fig. 1A), 33.1% of videos were deemed unusable due to hands being out of frame, poor visibility, or incorrect task execution. Overall, 56.6% showed some quality degradation (eg, low resolution, poor lighting, or camera distance), with only 43.4% meeting clinical-grade standards, consistent with a prior study finding of 48.1% of remote videos being low quality.5 We systematically reviewed the quality of the collected videos and compared them to “clinical-grade quality” (defined as videos with resolution ≥480px, frame rate ≥ 24 fps, adequate lighting, and proper framing) considered necessary to support accurate automated MDS-UPDRS scoring. While some degraded videos could potentially be interpretable by clinicians, they were unsuitable for reliable algorithm-based analysis. Based on the above definition, low resolution (≤480 × 360) was observed in 10.8% of videos, while 24.7% had frame rates below 24 fps—conditions that hinder the accurate detection of tremor and rapid finger movements. In 13.6% of recordings, participants’ hands were partially or completely out of frame, especially during movement tasks, often because of improper camera positioning. Furthermore, 9.1% of participants sat too far from or too close to the camera, affecting movement detection. Poor lighting which can affect effectiveness of landmark detection algorithms was seen in 25.1% of videos. Task execution errors, including recording only one hand when both were required, performing the wrong task, or stopping the recording prematurely, were noted in 18.8% of videos (Fig. 1B). Our experience suggests that routine remote video-based assessments in PD continue to face persistent challenges. To improve reliability, protocols should include enhanced instructional materials with task demonstrations and guided trial sessions, as well as real-time quality control, such as automated alerts for poor framing, lighting, or incomplete tasks. Studies identifying which quality issues most affect scoring accuracy could help refine feedback mechanisms. Establishing minimum technical standards (eg, resolution, frame rate) and allowing participants to review and re-record videos may further reduce unusable data and ensure robust, reliable automated assessments of Parkinson's disease motor performance. Sincerely, (1) Research project: A. Conception, B. Organization, C. Execution; (2) Data analysis: A. Design, B. Execution, C. Review and Critique; (3) Manuscript: A. Writing of the first draft, B. Review and critique. A.I.: 1B, 1C, 2A, 2C, 3B. S.M.: 2B, 3B. M.G.: 1C, 3B. K.W.P.: 1B, 1C, 3B. T.M.: 3B. H.D.: 1C, 3B. J.A.: 1B, 3B. M.S.M.: 1A, 1B, 1C, 2C, 3A. M.J.M.: 1A, 1B, 1C, 2C, 3B. We gratefully acknowledge the participants of the Canadian Open Parkinson Network (C-OPN) registry for their valuable contributions. Ethical Compliance Statement: The research project dataset was reviewed and approved by the University of British Columbia Clinical Research Ethics Board (Approval IDs: H18-03548 and H22-03748). Informed written consent was obtained from all participants prior to conducting interviews and video recordings. Participants were informed of their right to withdraw from the study at any time. All collected data are securely stored, remain confidential, and no personal identifiers, such as participants’ names, are included in any research reports. We confirm that we have read the Movement Disorders Journal's position on issues involved in ethical publication and affirm that this work is consistent with those guidelines. Funding Sources and Conflicts of Interest: This research was supported by a Collaborative Health Research Project grant from the Canadian Institutes of Health Research (CIHR) in collaboration with the Social Sciences and Humanities Research Council of Canada (SSHRC) and the Natural Sciences and Engineering Research Council of Canada (NSERC) [Grant No. GR013210], as well as by the Pacific Parkinson's Research Institute [Grant No. GR005879]. The authors declare that there are no conflicts of interest relevant to this work. Financial Disclosures for the previous 12 months: M.J.M. is supported by the John Nichol Chair in Parkinson's Research, Canadian Institutes of Health Research, and CHRP grant CPG-163986. All other authors report no financial disclosures. The raw video recordings of patients cannot be shared due to compliance with privacy regulations. However, the extracted features from these videos are available from the corresponding author upon reasonable request.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".