Scaling traffic variables from sensors sample to the entire city at high spatiotemporal resolution with machine learning: applications to the Paris megacity
Bibliographic record
Abstract
Abstract Road transportation accounts for up to 35% of carbon dioxide and 49% of nitrogen oxides emissions in the Paris region. However, estimates of city traffic patterns are often incomplete and of coarse spatio-temporal resolution, even where extensive networks of sensors exist. This study uses a machine learning approach to analyze data from 2086 magnetic road sensors across Paris, generating a detailed dataset of hourly traffic flow and road occupancy covering 6846 road segments from 2018 to 2022. Our model captures flow and occupancy with a symmetric mean absolute percentage error of 37% and 54% respectively, providing high-resolution insights into traffic patterns. These insights allow for the creation of a comprehensive map of hourly transportation patterns in Paris, offering a robust framework for assessing traffic variables for each significant road link in the city. The model’s ability to incorporate an emission factor based on the mean speed of the vehicle fleet, derived from flow and occupancy data, holds promise for developing a detailed CO 2 and pollutant inventory. This methodology is not limited to Paris; it can be applied to other urban centers with similar data availability, highlighting its potential as a versatile tool for sustainable urban monitoring.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".