Distributed Federated and Incremental Learning for Electric Vehicles Model Development in Kafka-ML
Bibliographic record
Abstract
With the increasing development and deployment of new systems for efficient and clean mobility, Electric Vehicles (EVs) are becoming more and more common among people. Those produce large amounts of data streams that need to be collected and analyzed to understand user needs and improve their performance. For this purpose, Artificial Intelligence (AI) techniques are playing a very important role. Within this context, Kafka-ML is a Machine Learning (ML) framework that enables the consumption and processing of data streams and allows the flexible management and deployment of neural networks throughout their entire life cycle. Kafka-ML can work with Distributed Neural Networks (DNN) which reduce latency and response times, perform incremental training over time allowing models to adapt to data on the fly, and carry out Federated Learning (FL) processes for this type of algorithms so a more robust global model can be created while maintaining data privacy and security, but all this separately. This work has considered the joint implementation of FL, for anonymous data sharing, incremental learning for continuous training of the models, and DNN for distribution of the models across different points on the map. All this applied within a Vehicle-to-everything (V2X) domain where EV usage and charge data can be shared to improve the user experience, as well as to better understand the behavior of this type of vehicles and their charging points to achieve savings, and how it affects people daily lives. An evaluation of the system related to this EV use case is presented to demonstrate the viability of the tool.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".