Data for: A novel dataset of indoor environmental conditions in work-from-home settings
Bibliographic record
Abstract
During the last week of March 2020, millions of workers worldwide transitioned to working from home due to the COVID-19 pandemic. The post-pandemic rate of remote work is likely to remain higher than pre-pandemic levels, making remote or hybrid work an intriguing area of study. IEQ research in work-from-home (WFH) settings is a relatively new field, with only a limited number of studies conducted so far. To address gaps in this research, a systematic field study was launched in 2022 to explore these environments. The study's objective was to assess both monitored and perceived IEQ, as well as the well-being and productivity of workers in home-based settings. The first paper from this project details the methodology used to collect IEQ and contextual data, publishes the collected data to the public domain, and provides a statistical analysis of the IEQ data in relation to various contextual factors, such as cooking habits and type of residence, which were also gathered during the study. The questionnaire-based contextual dataset will be published separately. This monitoring-based dataset comprises hourly measurements of IEQ indicators collected over a seasonal study campaign. Each participant was provided with an indoor desktop IEQ monitor (Awair Omni), along with installation instructions and protocols for placing these monitors in their WFH offices. The monitors recorded total volatile organic compounds (tVOC), particulate matter (PM2.5), carbon dioxide (CO2), air temperature, humidity, and sound pressure levels (SPL). The dataset is presented in an Excel spreadsheet format with separate worksheets for each of the IEQ indicators. The first column in the spreadsheet contains the date-time index, with each row representing an hour during the monitoring period. Each column represents a WFH site. The first 80 sites are from the Metro Vancouver area, the next 11 (site #81-91) from the Seattle Metropolitan area, and the last four sites (site #92-95) are located on Vancouver Island. The units for the monitoring data are as follows: tVOC (ppb), PM2.5 (μg/m³), CO2 (ppm), Temperature (°C), Relative Humidity (%), and SPL (dBA). To our knowledge, this study presents the first extensive dataset focused on WFH environments in Canada and is one of the few international studies to make collected IEQ data publicly available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.009 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".