Refining Sparse Cell-ID Trajectory of Public Service Vehicles by Spatiotemporal Modelling
Bibliographic record
Abstract
Mobile phone data have become a critical data source for transportation research. While a cell-id trajectory was routinely reorganized by International Mobile Subscriber Identity (IMSI), it potentially allows to analyze transportation behaviors and social interaction of total population, with a full temporal coverage at low cost. However, cell-id trajectory is often sparse due to low reporting frequency and uncertainness of mobile holders’ position. So, the cell-id trajectory refinement has been recognized as challenging work to further facilitate trajectory data mining. This paper presents a comprehensive approach to identify cell-id trajectories of public service vehicles (PSVs) from large volume of trajectories and further refines these cell-id trajectories by a heuristic global optimization approach. The modified longest common subsequence (LCSS) method is used to match a cell-id trajectory and a public transportation route (PTR) and correspondingly calculates their similarities for determining whether the trajectory is PSV mode or not. Taking full advantages of the nature of a PSV tends to move on the PTR in uniform motion to meet a prescript visit to stops, a heuristic global optimization approach is deployed to build a spatiotemporal model of a PSV motion, which estimates new locations of cell-id trajectories on the PTR. The approach was finally tested using Beijing cellular network signaling datasets. The precision of PSV trajectory detection is 90%, and the recall is 88%. Evaluated by our GNSS-logged trajectories, the mean absolute error (MAE) of refined PSV trajectories is 144.5 m and the standard deviation (St. Dev) is 81.8 m. It shows a significant improvement in comparison of traditional interpolation methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".