LoChain: A Decentralized and Privacy-Preserving Blockchain Protocol for Mobility Data Management
Bibliographic record
Abstract
Abstract. Mobility data has become a strategic asset in urban planning, crisis management and smart city operations. However, centralized systems for mobility tracking raise severe privacy concerns as they have the ability to directly link individuals to their movements. To address these issues, we propose LoChain, a decentralized protocol that enables the privacy-preserving collection and processing of mobility data based on blockchain technology. More precisely, LoChain replaces precise coordinates with standardized geoaddresses, associate user movements to disposable identities, communication them via Tor routing and stores the resulting data across a decentralized network built on Hyperledger Fabric. The system also employs a novel geopool and multi-channel architecture to simulate sharding, enabling localized data ingestion, inter-district communication and global statistical aggregation without compromising individual privacy. Localized position obfuscation and pseudo-random identity purging are used to further prevent reidentification. A proof-of-concept prototype, including an Android app, blockchain backend and visualization layer was developed and evaluated using synthetic data from 10,000 virtual users. The experiments results obtained from the simulation highlight the LoChain’s ability to preserve user privacy while maintaining analytical utility. Finally, we also introduce an incentive model as well as a decentralized governance structure to ensure long-term scalability, regulatory compliance and participatory control.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".