Bibliographic record
Abstract
Today's state-of-the-art machine learning (ML) techniques, such as deep learning (DL) networks are typically trained using cloud platforms, leveraging elastic scalability of the cloud. For such processing, data from various sources need to be transferred to a cloud server. While this works well for some application domains, it is not suitable for all applications due to concerns about latency, connectivity, and privacy. For example, sharing life logging photos and videos from cellphones and wearable devices can cause privacy concerns for users, and transferring the unstructured data can burden the communication network. With the increase of such applications, federated learning (FL) is proposed as a distributed ML solution for learning on edge devices, such as cellphones and wearable devices. In FL, clients collaboratively train a model on their device without sharing their data. Each client trains a local model with their data and shares the model parameters with a FL server to aggregate and build a global model. Shifting from traditional ML techniques to federated solutions requires comparing these two approaches. Moreover, users need to study the performance of FL models to decide if federation is feasible for their learning task. In this paper, we propose an automated solution to compare centrally trained DL models with federated solutions. The tool allows users to easily analyze the accuracy of federated models for their learning task and study the effect of the federated parameters. We show the features of our tool building central and federated DL models from an input model structure for recognizing images in the MNIST benchmark dataset.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.007 | 0.048 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".