Real-Time Riders: A First Look at User Interaction Data from the Back End of a Transit and Shared Mobility Smartphone App
Bibliographic record
Abstract
A fundamental component of transit planning is understanding passenger travel patterns. However, traditional data sources used to study transit travel have some noteworthy drawbacks. For example, manual collection of travel surveys can be expensive, and data sets from automated fare collection systems often include only one transit system and do not capture multimodal trips (e.g., access and egress mode). New data sources from smartphone applications offer the opportunity to study transit travel patterns across multiple metropolitan regions and transit operators at little to no cost. Moreover, some smartphone applications integrate other shared mobility services, such as bikesharing, carsharing, and ride-hailing, which can provide a multimodal perspective not easily captured in traditional data sets. The objective of this research was to take a first look at an emerging data source: back-end data from user interactions with a smartphone application. The specific data set used in this paper was from a widely used smartphone application called Transit that provides real-time information about public transit and shared mobility services. Visualizations of individuals’ interactions with the Transit app were created to demonstrate three unique aspects of this data set: the ability to capture multicity transit travel, the ability to capture multiagency transit travel, and the ability to capture multimodal travel, such as the use of bikeshare to access transit. This data set was then qualitatively compared with traditional transit data sources, including travel surveys and automated fare collection data. The findings suggest that the data set has potential advantages over traditional data sources and could help transit planners better understand how passengers travel.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".