Understanding User Behavior Through the Use of Unsupervised Anomaly Detection: Proof of Concept Using Internet of Things Smart Home Thermostat Data for Improving Public Health Surveillance
Bibliographic record
Abstract
BACKGROUND: One of the main concerns of public health surveillance is to preserve the physical and mental health of older adults while supporting their independence and privacy. On the other hand, to better assist those individuals with essential health care services in the event of an emergency, their regular activities should be monitored. Internet of Things (IoT) sensors may be employed to track the sequence of activities of individuals via ambient sensors, providing real-time insights on daily activity patterns and easy access to the data through the connected ecosystem. Previous surveys to identify the regular activity patterns of older adults were deficient in the limited number of participants, short period of activity tracking, and high reliance on predefined normal activity. OBJECTIVE: The objective of this study was to overcome the aforementioned challenges by performing a pilot study to evaluate the utilization of large-scale data from smart home thermostats that collect the motion status of individuals for every 5-minute interval over a long period of time. METHODS: From a large-scale dataset, we selected a group of 30 households who met the inclusion criteria (having at least 8 sensors, being connected to the system for at least 355 days in 2018, and having up to 4 occupants). The indoor activity patterns were captured through motion sensors. We used the unsupervised, time-based, deep neural-network architecture long short-term memory-variational autoencoder to identify the regular activity pattern for each household on 2 time scales: annual and weekday. The results were validated using 2019 records. The area under the curve as well as loss in 2018 were compatible with the 2019 schedule. Daily abnormal behaviors were identified based on deviation from the regular activity model. RESULTS: The utilization of this approach not only enabled us to identify the regular activity pattern for each household but also provided other insights by assessing sleep behavior using the sleep time and wake-up time. We could also compare the average time individuals spent at home for the different days of the week. From our study sample, there was a significant difference in the time individuals spent indoors during the weekend versus on weekdays. CONCLUSIONS: This approach could enhance individual health monitoring as well as public health surveillance. It provides a potentially nonobtrusive tool to assist public health officials and governments in policy development and emergency personnel in the event of an emergency by measuring indoor behavior while preserving privacy and using existing commercially available thermostat equipment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".