Refining Sparse Cell-ID Trajectory of Public Service Vehicles by Spatiotemporal Modelling
Bibliographic record
Abstract
Mobile phone data have become a critical data source for transportation research. While a cell-id trajectory was routinely reorganized by International Mobile Subscriber Identity (IMSI), it potentially allows to analyze transportation behaviors and social interaction of total population, with a full temporal coverage at low cost. However, cell-id trajectory is often sparse due to low reporting frequency and uncertainness of mobile holders’ position. So, the cell-id trajectory refinement has been recognized as challenging work to further facilitate trajectory data mining. This paper presents a comprehensive approach to identify cell-id trajectories of public service vehicles (PSVs) from large volume of trajectories and further refines these cell-id trajectories by a heuristic global optimization approach. The modified longest common subsequence (LCSS) method is used to match a cell-id trajectory and a public transportation route (PTR) and correspondingly calculates their similarities for determining whether the trajectory is PSV mode or not. Taking full advantages of the nature of a PSV tends to move on the PTR in uniform motion to meet a prescript visit to stops, a heuristic global optimization approach is deployed to build a spatiotemporal model of a PSV motion, which estimates new locations of cell-id trajectories on the PTR. The approach was finally tested using Beijing cellular network signaling datasets. The precision of PSV trajectory detection is 90%, and the recall is 88%. Evaluated by our GNSS-logged trajectories, the mean absolute error (MAE) of refined PSV trajectories is 144.5 m and the standard deviation (St. Dev) is 81.8 m. It shows a significant improvement in comparison of traditional interpolation methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".