Estimating Pedestrian Volumes at Intersections using Artificial Intelligence: A ChatGPT Vision Approach
Bibliographic record
Abstract
Pedestrian volume at intersections is a key input to transportation processes and is vital for developing pedestrian-centric interventions. Though technological solutions to provide continuous monitoring of intersections and field volume counts do exist, most jurisdictions do not have this technology widely deployed and therefore still require a means of estimating volumes. Conventional estimation methods, such as expansion factor methods or the development of direct-demand (DD) models, are alternatives typically employed by practitioners, but they still require known pedestrian volume data for numerous sites, which are often not readily available for jurisdictions. This work explores the use of ChatGPT-4 Vision (GPT-4V) as a tool to assist jurisdictions in estimating pedestrian volume within intersections. The initial assessment of GPT-4V demonstrated its capability in interpreting satellite images, particularly in ranking sites according to pedestrian activity levels. Subsequently, a method was implemented to rank 48 sites based on pedestrian activity using GPT-4V and satellite images. A linear correlation of 0.73 was achieved between the GPT-4V ranking and the true ranking determined from observed field data. Following that, a method was proposed to combine the GPT-4V site rankings and field volumes collected at selected key intersections to estimate the pedestrian volume at all the sites ranked using GPT-4V. Its performance rivaled (and sometimes surpassed) that of existing conventional methods like DD models, all without the need for complex statistical models or extensive datasets. This simplicity makes the GPT-4V method promising for cost-effective pedestrian exposure estimation. This work also investigates inconsistencies and biases in GPT-4V’s responses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".