Classification of natural indoor and outdoor scenes from radiometric, photometric, and colorimetric features
Notice bibliographique
Résumé
Background: In natural, real-world scenes, environmental light can vary significantly in intensity, spectrum, color, and spatial and temporal characteristics due to the presence of different light sources (daylight, electric lighting, self-luminous displays, and mixtures thereof) and light-surface interactions. The ‘spectral diet’ an observer is exposed to reflects this complexity, compounded by body, head, and eye movements. Understanding some of this complexity requires a systematic and detailed understanding of the variation of light in the real world. From the first principles, we know that light outdoors is of higher intensity than light indoors. A detailed survey of the spectral, spatial, and temporal features of light in the real world, which is highly relevant for architectural design, occupational health, and environmental medicine, has thus far not been undertaken. Methods: In the SCENES Dataset (https://www.scenes-dataset.org/), we have comprehensively characterized the spectral, spatial, and temporal variations of natural scenes using a novel multimodal data collection setup comprising an α- opic imaging radiometer, a high-resolution spectroradiometer, illuminance and colorimetric measurement devices, a depth camera, and an uncalibrated wide- field RGB video camera. All instruments were integrated into a portable box for easy deployment with a power supply through external batteries. Data were collected across various times of day and seasons. Each scene was described using a novel metadata descriptor (n=43 items), encompassing detailed information, including geographical information, weather conditions, and scene categories (9 indoor subcategories, 11 outdoor subcategories). To understand the basic aspects of the datasets, we used descriptive statistics (mean, SD, min., max.), and applied the random forest algorithm to develop a scheme for indoor vs. outdoor specifications on spectrum-derived data. Results: We measured natural scenes both indoors (n=313) and outdoors (n=366) in five locations (Tübingen, Germany: n=53 indoors, n=254 outdoors; Munich, Germany: n=61 indoors, n=11 outdoors; Prague, Czech Republic: n=19 indoors, n=16 outdoors; Lyon, France: n=67 indoors, n=26 outdoors; Ottawa, Canada: n=113 indoors, n=59 outdoors). As expected, there are key differences between indoor vs. outdoor scenes in photopic illuminance (mean±SD 1826.51±7402.82 lx [min. 13.8 lx, max. 63721.1 lx] indoors vs. 13401.91±7402.82 lx [0.28 lx, 110872.8 lx] outdoors) and melanopic equivalent daylight illuminance (mean±SD 1586±6416.93 lx [min. 11.9 lx, max. 55178.3 lx] indoors vs. 12240.01±18741.81 lx [0.19 lx, 97807.52 lx] outdoors). The results from the random forest algorithm indicate excellent model performance (accuracy: 0.963, precision: 0.982, recall: 0.931, F1: 0.956). Regarding feature importance in the classification, CRI Ra (color rendering index) was ranked highest. Conclusion: In this study, we collected indoor and outdoor natural scene data across various geographical contexts. A preliminary analysis has evaluated the ability of several spectrum-derived features to allow for the classification of indoor vs. outdoor scenes. Further analyses will investigate the possibility of classifying scene subcategories and probe the spatiotemporal characteristics of natural scenes in further detail. The SCENES Dataset will be made available as an open-access dataset to serve as a benchmark of the properties of environmental light.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,006 | 0,005 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,001 | 0,002 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».