Prototype Development of a Responsive Emotive Sensing System (DRESS): System Operations Testing Outcomes
Bibliographic record
Abstract
Background: Smart-home consumer technologies have been criticized for failing to disclose their operational performance characteristics to the marketplace. As one result, some users of wearable fitness technologies have reported being frustrated by invalid motivational responses based on fluctuations in accurate performance measurement by certain brands. Gerontechnology researchers have similarly documented the critical importance of valid operations and technical stability as major influences on whether older adults and their caregivers adopt and use new cognitive assistive technologies. We have been iteratively developing the DRESS (Development of a Responsive Emotive Sensing System) system, integrating context aware computing with effective sensor and interactive technologies, to customize coaching persons with dementia to dress independently. Our prior testing focused on components and clothing identification, not the overall system performance. Consequently, we initiated system testing, as part of our alpha version development phase, to assess key metrics and disclose the performance outcomes. Objective: To assess the operational accuracy (validity) and stability (reliability) of the DRESS system alpha prototype model. Methods: We conducted a 110 day device trial run-in study. The system operated 24/7 in a studio-sized testing unit using the local WiFi network. A 69-year-old tester documented any usability issues during this period. Automatic log reports were generated daily by the system and validated and annotated by the project manager. A content analysis of the user and log reports was conducted, and descriptive statistics were used to describe the operational findings. Results: The system functioned error free for the majority of the trial (75% of days) with stable performance for 95.5% of days. Thirty-seven correctable error events occurred during 28 of the 110 days and resulted in 4 categories of errors: Hardware (0.9%), from a defective IPad charger; Network (3.6%), from host network disconnects/power outage; Usability (4.5%), from the visual displays/buttons on the caregivers’ device being too small in size; and Re-initialization (24.5%), from the operating system/Indigo software updates. Conclusions: Overall, the system performed very favorably for an alpha prototype. As expected, the initial deployment required an immediate debugging period primarily rectified by software recoding. Notably no fatal or irresolvable errors occurred. The system remained stable except for a disconnect due to a weather-related regional power outage. Lessons learned, such as integrating a remote automatic reboot capability, will be used to further optimize system performance before advancing to an in-home study with persons experiencing dementia.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".