Recent Advances in Complexity Reduction Methods for VVC Inter Coding: A Review
Bibliographic record
Abstract
Video traffic continues to surge, pushing codecs toward higher efficiency at the cost of sharply increased complexity. The H.266/Versatile Video Coding (VVC) standard roughly halves bitrate relative to HEVC at comparable quality but increases encoder complexity substantially. This paper surveys algorithm-level complexity-reduction methods for VVC inter-frame coding, grouping them by decision module—coding unit (CU) partitioning, inter-mode selection, and motion estimation (ME)—and by approach—heuristic/statistical, machine learning (ML), and deep learning (DL). We provide quantitative comparisons of reported complexity–efficiency trade-offs (encoder time vs. BD-rate), and discuss observed trends and potential areas for improvement. Across the literature, CU partitioning is the most heavily targeted module, followed by ME; reported encoder time savings span approximately 10–55% with typically no more than about 3% BD-rate increase, depending on module and method. It is observed that DL methods generally achieve the largest time savings (often around 50% or higher) by predicting partition structures or mode decisions end-to-end, at the expense of training data and inference cost. Heuristic methods remain lightweight and hardware-friendly with small BD-rate impact, and ML methods provide balanced trade-offs. To the best of our knowledge, this is the first comprehensive survey of VVC inter-frame coding, distilling practical lessons and outlining open directions—especially hybrid pipelines that combine inexpensive filters with learned predictors—to guide more efficient VVC implementations and future standards.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".