MétaCan
Menu
Back to cohort
Record W4386053118 · doi:10.1109/mdm58254.2023.00033

Trajectory-User Linking using Higher-order Mobility Flow Representations

2023· article· en· W4386053118 on OpenAlexaff
Mahmoud Alsaeed, Ameeta Agrawal, Manos Papagelis

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldSocial Sciences
TopicHuman Mobility and Location-Based Analysis
Canadian institutionsYork University
Fundersnot available
KeywordsComputer scienceTrajectoryEmbeddingArtificial intelligenceEncoderMachine learningRepresentation (politics)SegmentationData miningData modelingTheoretical computer science

Abstract

fetched live from OpenAlex

Trajectory user linking (TUL) is a problem in trajectory classification that links anonymous trajectories to the users who generated them. TUL has various uses such as identity verification, personalized recommendation, epidemiological monitoring, and threat assessments. A major challenge in TUL modeling is sparse data. Previous TUL research heavily relies on sequence-to-sequence models such as RNNs and LSTMs, with trajectory segmentation to combat sparsity, but segmentation does not sufficiently address the issue and existing models often ignore data skewness, resulting in poor precision and performance. To address these problems, we present TULHOR, a TUL model inspired by BERT, a popular language representation model. One of TULHOR’s innovations is the use of higher-order mobility flow data representations enabled by geographic area tessellation. This allows the model to alleviate the sparsity problem and also to generalize better. TULHOR consists of a spatial embedding layer, a spatial-temporal embedding layer and an encoder layer, which encodes properties and learns a rich trajectory representation. It is trained in two steps, first using a masked language modeling task to learn general embeddings, then fine-tuned using a balanced cross-entropy loss to make predictions while handling imbalanced data. Experiments on real-life mobility data show TULHOR’s effectiveness as compared to current state-of-the-art models.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.179
Threshold uncertainty score0.996

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.002
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0050.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.068
GPT teacher head0.373
Teacher spread0.305 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes1
Has abstractyes

Explore more

Same topicHuman Mobility and Location-Based AnalysisFrench-language works237,207