Outils de visualisation de données de cartes à puce pour une société de transport collectif
Bibliographic record
Abstract
Public transit authorities are choosing more and more smart card automated fare collection systems and realize that those daily recovered data, since 2008 for the greater Montreal region, have a great potential for their planning and operations.In this context, this research master is part of a global project held for three-year period in collaboration with various partners.It follows previous research works on data enrichment of smart card transactions by combining their trip origin and destination.For the purpose of this project, the transit authority RTL (Rseau de Transport de Longueuil) provided one month (March 2013) of bus and metro smart card transactions (3.1 million).As far as Thales is concerned, they made available their "Analytics For Transportation" portal developed by its CeNTAI Department (Centre de Traitement et d'Analyse de l'Information).The main objective of this master research is to design interfaces for viewing and analyzing smart card transactions, enriched of their destination, while meeting the needs of a transit operator.The sub-objectives, corresponding to the steps of this research, are:-Make operational the algorithm determining trip destinations -Conceptualize the most adequate data structure enabling their visualization -Design visualization interfaces meeting the needs of a transit operator This thesis starts with a literature review with, on the one hand, the previous works on the estimation of the trips origin and destination, and, on the other hand, other projects on data visualization.The steps followed to meet the above three sub-objectives are described in the methodology section.The final section presents the results and analysis obtained from these enriched data.The main achievements of this project are:-The optimization and redesign of the algorithm estimating trip destinations and its adaptation to a network defined with the GTFS format (General Transit Feed Specification)The presentation of ergonomic insights, obtained thanks to the use open source tools (Elasticsearch, Kibana), enabling those enriched smart card data to be quickly analyzed viii -The design of a new customized web interface developed to present other key indicators used by a public transport companyIn conclusion, this research project presents an operational solution, which for a set of smart card transaction data offers, in one step, to estimate the destination of each smart card transaction trip, to prepare additional statistics (distance and travel time, trip-leg sequences ) and to export those enriched transactions to a text file or a data base (Elasticsearch).The whole process is made within a relatively short time: 20 minutes for 3 million transactions, export time included.The data is then directly available and usable in web portals configured or developed for the occasion and which take into account the needs of the customers.Of the 3.1 million available transactions, 20% are metro transactions.These transactions help the algorithm in the estimation of a trip destination.These metro transactions only help to find 1 more percent of destinations, resulting in 79% of trip destinations recovered for our March 2013 dataset.Trip-legs have also been reconstructed by the algorithm.It shows for example that 66% of bus travels are made without a transfer.The share of users making only one transfer represents respectively 12% from bus to bus and represents 20% from bus to metro.In the end, this research shows that the analysis of large volume of data within a limited period of time is possible and an operational solution is presented.Indeed, it would require a processing time of 32 hours to enhance the RTL smart card transactions of the last 8 years, with 3 million transactions per month.These OD type of data would then be available to power the analysis of the various departments of a public transit authority such as operations, planning and even marketing and finance.The developed visualization prototypes would then help the RTL in drafting the specifications of a new tool sold and designed by a company selling BI (Business Intelligence) solutions to visualize their business data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".