Learn to Compress (LtC): Efficient Learning-based Streaming Video Analytics
Bibliographic record
Abstract
Video analytics are often performed as cloud services in edge settings, primarily to offload computation and also in situations where the results are not directly consumed at the video source. Sending high-quality video data from end devices can be expensive in terms of both bandwidth and power use. To build a streaming video analytics pipeline that makes efficient use of these resources, it is imperative to reduce the size of the video streams. Traditional video compression algorithms are unaware of the semantics of the video, and can be both inefficient and harmful to the analytics performance. In this paper, we introduce LtC, a collaborative framework between the video source and the analytics server that efficiently learns to reduce the video streams within an analytics pipeline. Specifically, LtC uses the full-size video analytics algorithm at the server as a teacher to train a lightweight student neural network, which is then deployed at the video source. The student network is trained to capture the semantic significance of different regions within a video, which is used to selectively preserve the crucial regions in high quality while aggressively compressing the remaining regions. Furthermore, LtC incorporates a novel temporal filtering algorithm based on feature differencing to omit transmitting frames that do not contribute new information. Overall, LtC reduces bandwidth usage by 28-35% and attains a response delay that is up to 45% shorter than current state-of-the-art methods, while maintaining comparable analytics performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".