Activity monitoring system using deep learning for people with dementia
Bibliographic record
Abstract
Dementia is a degenerative condition that affects cognitive abilities and daily functioning. This project aims to explore and evaluate activity recognition algorithms to support the assisted living of people with dementia. The proposed deep learning approach can help to monitor people with dementia and support their caregivers in providing effective care. We tried a new approach for detecting the activities of daily living for people with dementia. We explored ExpansionNet_v2 model and used it to train on the Toyota Smart Home dataset in order to detect the activities od daily living. The dataset was converted into COCO dataset format. Bounding boxes were generated using Faster-RCNN with ResNet backbone pretrained model from pytorch. Captions were generated using scene understanding. This involved analyzing the image or video to extract semantic information about the environment and objects within it, including their relationships and context. Semantic relationships and patterns were extracted, which helped in building a more comprehensive understanding of the scene. The training process involved two steps - initial training and fine-tuning. During initial training, newly added layers were trained while keeping the pre-trained layers of the Swin-Transformer backbone frozen. Fine-tuning involved training the entire network, including both the pre-trained backbone and newly added layers, on the dataset. The purpose of using multiple frames from a video during training is to increase the probability of detecting the pose accurately and generating a good caption. The algorithm to detect ADL was tested on real-life videos of three dementia patients at different stages of dementia. The daily activities of these patients were recorded to test the algorithm after training and validation on the Toyota SmartHome dataset.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".