Drawing Heritage Façades from 3D Data: Reprocessing LiDAR and Photogrammetry Datasets from OpenHeritage3D
Bibliographic record
Abstract
Façade documentation requires detailed drawings that illustrate the dimensions, locations, conditions, and materials of a building exterior. This documentation aims to provide a guide and record for assessing, monitoring, preserving, and restoring built heritage. In 1858, a Prussian architect, Albrecht Meydenbauer, nearly fell while taking direct measurements at a building façade. This incident motivated him to develop an indirect measurement method called photogrammetry to record built heritage façades using a photographic camera and surveying instruments. In the era of analogue photogrammetry, photographic films and other equipment were costly. Consequently, such photogrammetric documentation projects required careful planning and skilled personnel. In the 21st century, the cost of personal computers and digital cameras reduced significantly. Collecting large amounts of data has become relatively inexpensive. Meticulous planning in photogrammetric projects has been de-emphasized, and tools like digital cameras, LiDAR scans, and UAVs have significantly increased the volume of unstructured data collected. CyArk, a non-profit organization based in California, has documented heritage sites globally using a combination of LiDAR and photogrammetry and published their datasets through the OpenHeritage3D portal, an open-access platform. However, these datasets often lack proper description, documentation, categorization, standardization, and modularization. Although these data can create photorealistic models, they do not contain surveying data and are too large to process for most personal computers. This project proposes a workflow for reprocessing these LiDAR and photogrammetric datasets produced by CyArk using widely available software like RealityCapture, AutoCAD, AutoCAD Raster Design, and the Bulk Rename Utility. The objective is to deliver traditional 2D façade documentation from LiDAR and photogrammetric data required for heritage conservation projects and to create guidelines for the future collection and storage of such large building documentation datasets. Preliminary results show that more emphasis should be placed on metadata, georeferencing, and data organization than has hitherto been made.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.009 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.011 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".