Imagery datasets for photobiological lighting analysis of architectural models with shading panels
Bibliographic record
Abstract
This paper describes eight imagery datasets including around 12000 images grouped in 1220 sets. The images were captured inside an architectural model aimed at exploring the impact of shading panels on photobiological lighting parameters. The architectural model represents a generic space at 1:10 scale with a single side fully glazing façade used to install shading panels. The datasets present interior lighting conditions under different shading configurations in terms of surface colors and glossiness, horizontal and vertical orientations and upwards, downwards, and left/right inclinations of panels, V-shape opening, low to high densities, and top and bottom positions at the window. The experiments of shading panel configurations were conducted under four to six different exterior overcast daylighting conditions simulated with very cool to very warm color temperatures and high to low intensities inside an artificial sky chamber. The datasets include bracketed low dynamic range (LDR) images which enable generating high dynamic range (HDR) images for photobiological lighting evaluations. Images were captured from the side and back viewpoints inside the model by using Raspberry Pi camera modules mounted with fisheye lenses. The datasets are reusable and useful for architects, lighting designers, and building engineers to study the impact of architectural variables and shading panels on photobiological lighting conditions in space. The datasets will also be interesting for computer vision specialists to run machine learning techniques and train artificial intelligence for architectural applications. The datasets are partially used in Parsaee, et al. [1]. The datasets are compiled as part of a doctoral dissertation in architecture at Laval University authored by Mojtaba Parsaee [2]. The datasets are shared through two Mendeley data repositories [3,4].
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".