I-FCSAM: An integrated framework of few-shot learning and segment anything model for vision-based indoor built environment management
Bibliographic record
Abstract
Accurate and timely analysis of as-is versus as-planned conditions is critical for built environment management (BEM) in the Architectural, Engineering, Construction, and Operation (AECO) sector. Various AECO applications, such as progress monitoring, facility management, quality control, and inspections, rely on comparative analysis for effective decision-making. However, the manual analysis can be time-consuming and prone to errors. Computer vision methods offer promising solutions; however, their adoption in the AECO face challenges due to high data annotation requirement, computational demands, and limited datasets. This challenge is further augmented in indoor built environment management (IBEM) compared to outdoor environments due to more object diversity, data scarcity, and inadequate representation of AECO-specific objects in existing datasets. This necessitates implementation of vision-based systems with minimal training data and effort. Therefore, this study introduces the I-FCSAM framework, an integrated approach that combines Few-Shot Learning (FSL) and Segment Anything Model (SAM) to identify objects in indoor built environment visualizations with the ability to handle limited image samples available. The FSL model, based on Prototypical Networks (PNs), is implemented on 25 classes of AECO-specific objects, representing the as-is states of the indoor built environments. The integration of SAM’s segmentation with FSL’s classification capabilities enabled instance segmentation of the objects and minimized clutters. With FSL’s overall accuracy of 86.78%, the I-FCSAM framework demonstrated promising performance (78% precision and 61.1% recall) in classification of SAM-generated objects/regions, reducing the need for extensive labeled data compared to baselines and holding great potential for enhancing vision-based comparative analysis in IBEM applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".