Development of Indicators to Measure Transit Use Variability Based on Smart Card Data
Bibliographic record
Abstract
Transit ridership varies over time, within the behaviour of a same user, and from one user to another. These variations make it difficult to adjust service and demand forecasting models, potentially leading to additional operating costs and a non-optimal allocation of vehicles on the network. However, the growing availability of longitudinal and individualized data now allows to better understand travel behaviour. In particular, this paper benefits from smart card data from the OPUS system of Montreal, Canada. More than 429 million validations, made by nearly 2 million cards, are mined to investigate transit use variability at a totally disaggregated level over a one-year period. Four indicators are proposed to measure several types of variations: trip dispersion among the users, variability of the frequency of use, temporal variance of the monthly number of trips and spatial diversity of the boarding locations. These indicators are applied to evaluate the regularity of 10 distinct groups of cards defined by their fare composition. Because of the sensitivity of statistical tests on sample size, an effect size analysis is provided to better quantify the magnitude of the differences between the groups. The results reveal relationships between transit use and the typeof product used. Both annual and monthly pass users are found to be regular and frequent passengers, whereas ticket book users are rather occasional travellers. Moreover, variability tends to increase with the number of different products used during the year. L’achalandage du transport en commun varie dans le temps, au sein du comportement d’un même usager, mais aussi d’un usager à l’autre. Ces variations rendent difficile l’ajustement des services et des modèles de prévision de la demande, conduisant potentiellement à des coûts d’opération supplémentaires et à une affectation non optimale des véhicules sur le réseau. Toutefois, la disponibilité croissante de données longitudinales et individualisées permet désormais de mieux comprendre la variabilité des comportements de déplacement. Cet article bénéficie ainsi de données de cartes à puce provenant du système OPUS de Montréal, Canada. Plus de 429 millions de validations, réalisées par près de 2 millions de cartes, sont exploitées pour étudier la variabilité d’utilisation du transport en commun à un niveau totalement désagrégé sur une période d’un an.Quatre indicateurs sont proposés afin de mesurer plusieurs types de variations : la dispersion des déplacements parmi les usagers, la variabilité de la fréquence d’utilisation, la variance temporelle du nombre de déplacements par mois et la diversité spatiale des lieux d’embarquement. Ces indicateurs sont appliqués pour comparer la régularité de dix groupes de cartes distincts définis en fonction de leur composition tarifaire. Face aux limites des tests statistiques traditionnels, sensibles à la taille de l’échantillon, la notion de taille d’effet est introduite pour mieux quantifier l’importance des différences observées entre les groupes. Les résultats révèlent qu’il existe une relation entre l’utilisation du transport en commun et le type de titre utilisé. Les utilisateurs d’abonnements annuels ou mensuels sont en moyenne très réguliers et fréquents, alors que les utilisateurs de carnets de tickets sont des usagers plus occasionnels. De plus, la variabilité d’usage tend à augmenter avec le nombre de titres différents utilisés durant l’année.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".