Textual Data and Natural Language Processing for Next Generation Intelligent Transportation Systems: Perspectives, Techniques and Challenges
Bibliographic record
Abstract
RÉSUMÉ: RÉSUMÉ Les Systèmes de Transport Intelligents (STI) contribuent à l’amélioration des différents aspects des réseaux de transport tels que la sécurité, la fiabilité, les choix de voyage éclairés, la performance environnementale et la résilience de l’exploitation du réseau. Un changement significatif dans STI au cours des dernières années est que ces systèmes passent des systèmes traditionnels axés sur la technologie aux systèmes axés sur les données. Cependant, au milieu de ce changement, les données textuelles sont une sorte de données qui a été moins explorée. Depuis ces dernières années, beaucoup plus de données publiques sont désormais disponibles à partir de diverses sources de données qui peuvent être traitées pour aider à améliorer une variété de domaines dans les systèmes de transport. Les données des réseaux sociaux sont une forme spéciale de contenu textuel riche, généré par les utilisateurs, qui recèlent d’énormes potentiels d’utilisation dans divers domaines du transport. Cette thèse se concentre sur l’exploration des domaines dans lesquels les données textuelles, notamment les données des réseaux sociaux, peuvent être bénéfiques pour STI et propose des approches de traitement automatique du langage naturel (TALN) pour examiner le potentiel de ces données pour aider à améliorer plusieurs domaines des réseaux de transport. Cette thèse propose une séquence de méthodes nécessaires pour transformer les données textuelles en contenu utilisable dans la modélisation pour la pratique et la recherche en introduisant les concepts fondamentaux et en proposant les méthodes de métriques de similarité, de text embeddings, d’analyse de sentiment et de modélisation de topics, chacune abordant un do-maine de problèmes rencontrés dans l’application des techniques de TALN dans le domaine des transports. Pour mener cette étude, un plan de recherche a été formulé incluant des méthodes spécifiques, en quatre grandes parties, dont chacune implique une approche distincte. Toutes les parties sont liées en termes d’objectifs et diffèrent dans leur approche et leurs détails. La portée de l’étude a été limitée pour inclure uniquement le réseau du métro et comme études de cas, les réseaux de métro canadiens. Le premier article de cette thèse implique une approche de classification de texte multilingue pour détecter les incidents dans le réseau de transport du métro dans la région métropolitaine de Montréal au Québec, Canada. ABSTRACT: ABSTRACT Intelligent Transportation System (ITS) refers to technologies applied to transportation sys-tems that improve transportation outcomes such as safety, reliability, informed travel choices, environmental performance, and network operation resilience. A momentous change in ITS in recent years is that much more public data is now available from various data sources that can be processed to address a variety of areas in transportation. As the systems shift from traditional technology-driven systems to data-driven intelligent transportation systems, it leads to evolution in ITS development. Notwithstanding, in the midst of this change, textual data is one sort of data that has been less explored for transportation application. Social media data is a special form of rich textual content, generated by users, holding tremendous potentials to be used in a variety of areas in transportation. This thesis focuses on exploring the capacity of social media data where it can be beneficial to ITSs and propose natural language processing approaches to examine the potential usages of such data in transportation domain. To carry out this work, a research design was formulated including specific methods, each to address one area in the transportation field. The scope of the study was limited to include only Metro network and as case studies, Canadian subway network and the Montreal metro network. The study was conducted in four major parts, of which each involved a distinctive approach. All parts are related in terms of objectivity and differed in approach and detail. Part one involved an approach for multilingual text classification to alert incidents in the metro transportation network within the geolocation boundaries of Montreal metropolitan area in Quebec, Canada. It deals with classification methods combined with multiple text representation algorithms in order to perform incident detection from tweets. Empirical experiments in this study demonstrates the ability of Bidirectional Encoder Representations from Transformers approach to generalize to cross-lingual and multilingual representation for metro incident related tweet classification. Part two elaborates a methodology to assess passengers’ perceptions of transit services and the impact of incidents on transportation users based on social media data. It focuses on investigating the effectiveness of the usage of social media data in transportation studies to model and evaluate public transport riders’ sentiment in different situations and regarding different topics. The case study is conducted in Société de transport de Montréal STM metro network in Montreal, Quebec, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.007 | 0.011 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".