Optimizing the Data Processing Decision for Hybrid Fog Environments
Bibliographic record
Abstract
In the Cloud, computing loads found a new virtualized and dynamically scaling home. However, because of the pay-as-you-go business concept that is dominant in the Cloud, scenarios where a big percentage of requests turn up with low-value data (as in incomplete, meaningless or out-of-sync) can be financially detrimental to the Cloud tenant. Internet of Things offerings expand the list of data sources for a Cloud-based service. They also take the data filtering and pre-processing required of services to get to the core value-returning requests to a new level. The availability of Fog nodes offers an opportunity where some processing is done on the edge of the network in order to distribute the load and minimize the network congestion caused by low-value data. This poses a question to Cloud service designers on how to optimize this process. The decision as to where to perform each step of the data management can make the difference for Cloud providers in mitigating both the risk of pushing high loads to the Cloud servers and network which increases the cost and the risk of almost localizing the whole process and losing the benefits from Cloud services. This process is complicated by constraints pertaining to the Fog nodes capacity, and bandwidth available. Furthermore, the impact of realistic factors like Fog node ownership and device priority must be considered. To tackle this challenge, we consider the question of optimizing the decision process for a data-intensive highly distributed Cloud service. A novel optimization model is presented with 3 alternate objects (request computational delay, service provider cost and a weighted multi objective version). Initial experimental results using 4 heuristic algorithms are presented. Shown results offer some insight into the contradicting factors in play (cost, network capacity, node and Cloud capacity).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".