Enhancing the Spatial Transferability of Direct Demand Models for Estimating Pedestrian Volumes at Intersections
Bibliographic record
Abstract
Direct demand (DD) models are an important tool for estimating annual average daily pedestrian traffic (AADPT) for all intersections in a jurisdiction. These models associate socioeconomic and land-use variables with pedestrian exposure and allow the estimation of AADPT for sites where pedestrian counts are not (readily) available. However, some jurisdictions lack pedestrian volume counts from a sufficiently large number of intersections to develop their own DD model or do not have the institutional resources to carry out the model development. Under these circumstances, a cost-effective alternative is to use DD models that were developed in other jurisdictions. Previous research evaluated the spatial transferability of DD models in scenarios where no pedestrian counts are available (i.e., naïve transferability) and showed that this resulted in large estimation errors. This paper examines methods to improve the estimation accuracy of spatially transferred DD models by using AADPT that is readily accessible to jurisdictions (we call this local calibration). Five local calibration models were proposed and evaluated using observed field counts and synthesized counts from three jurisdictions. The best model to use is a function of the number of local jurisdiction sites for which pedestrian counts are available. When pedestrian volume is available for 10% of the sites, Model C presented the best results for the synthetic approach: an average improvement of 8.7% when comparing the locally calibrated and naïve estimates. Using real AADPTs and very limited samples for local calibration, Model C also presented the best performance: an average improvement of 35.0%.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".