A Method for Studying Traffic Congestion Using New Data: Focusing on the Canadian Regions of Toronto and Hamilton
Bibliographic record
Abstract
Traffic congestion plays an important role in shaping the public debate over transportation policies and programs. But although there are many studies of the intensity and causes of traffic congestion, there are opportunities to improve methods and congestion metrics using newly available traffic data. Common regionally-scaled congestion studies simplify the understanding of congestion. Better congestion metrics would both reflect geographic and temporal patterns of congestion and would identify underlying predictors in an effort to better manage gridlock. While such metrics have previously been intractable due to lack of appropriate data, recent improvements in private sector traffic data now make more complete understandings of congestion possible. This study uses a novel data source from Inrix, Inc. to design and test new modeling approaches to characterize and identify the predictors of urban congestion on the arterial network in the Toronto and Hamilton Census Metropolitan Areas (CMAs). Methods are developed which make two novel contributions. First, models explore the geography of congestion at different scales. Second, models incorporate a discrete-continuous modeling framework which better reflects congestion’s non-linear nature. When tested in the two study regions, results suggest that while congestion is largely a regional phenomenon in Toronto, it is highly localized in Hamilton. Different policy responses and roles for smart growth planning may be in order in each of the regions. This methodology can be deployed in other regions to similarly measure congestion and identify more context-appropriate responses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.011 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".