A Hybrid Regression Model for Video Popularity-Based Cache Replacement in Content Delivery Networks
Bibliographic record
Abstract
Content Delivery Networks (CDN) and their globally dispersed caches host a myriad of User Generated Videos (UGV) to meet end-user requests with quality of service. To efficiently utilize the limited storage of the caches, it is imperative to improve the hit ratio of UGVs. In contrast to the traditional static content, UGV popularity is highly dynamic and dependent on end-user behavior. Therefore, we devise a novel popularity prediction model for UGV, using a hybrid regression model. Our hybrid regression model dynamically adapts the popularity of UGV that is built from a historical training dataset. We reduce error in predicting popularity by up to 14%, when compared to pure offline and online approaches, with a small increase in the execution time and memory overhead. Our novel popularity prediction model accounts for end- user behavior by considering the end-user video watch time and the number of shares for the UGVs. To improve cache performance in CDN, we employ a cache replacement strategy that leverages our popularity prediction model to efficiently evict the less popular UGVs for more popular content. We compare our novel cache replacement strategy with the traditional and state-of-the-art cache replacement strategies and show an increase in the average hit ratio of up to 74% and 7%, respectively, for UGVs with shortterm popularity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".