What We Should Know About Housing Reconstruction Costs
Bibliographic record
Abstract
This paper presents two important analyses, which have been derived from a rich dataset supplied by the Toronto Dominion Insurance (TDI) company. These analyses provide us with an improved understanding of house values, via a detailed analysis of their predicted reconstruction costs. In an initial step, we propose a new model that focuses on the modeling of house reconstruction cost (HRC) using 16 predictors, which include the materials and composition of the buildings. In a second step, we analyze the distribution of HRC using quantile regressions, in order to gain a better understanding of the influence of HRC skewness, which is driven by the most expensive houses. It is found that when a broad set of (16) predictors is used, the Living Space alone accounts for 54.87% of the cost variation, while the square of this variable accounts for another 9.4% of the variation in cost. Quantile analysis provides additional information, as the impact of certain coefficients on the cost of less expensive houses is different to that of expensive houses. In particular, the Age of Construction coefficient at the 25th quantile is 3 times higher (in absolute value) than at the 99th quantile, whereas the quantiles estimates for the nonlinear influence increase with quantile houses. The Living Space predictor reveals that the living-space cost is approximately 4 times greater at the 25th than at the 95th quantile, whereas the nonlinear influence of living space varies from a negative effect (lower quantiles) to a positive effect (upper quantiles), suggesting that the cost curve changes from concave at lower quantiles to convex at higher quantiles.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.063 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.011 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.018 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".