A Deep Learning-Based No-Reference Quality Metric for High-Definition Images Compressed With HEVC
Bibliographic record
Abstract
An accurate no-reference image quality assessment metric for compression artifacts is essential for the broadcasting and streaming industries. Although we have witnessed impressive advances in the capturing, delivery and display technologies, we have not managed to match them with an accurate and perceptual based no-reference image quality metric. In this paper, we propose a unique perceptual based no-reference quality metric for compressed HD frames/images that is based on the DenseNet network architecture. We focus on the effect HEVC (High Efficiency Video Coding) compression artifacts have on the visual quality of a broadcasted and streamed video, as this is a requirement of immense importance for these industries. We chose the Video Multi-Method Assessment Fusion (VMAF) metric as our base measure to map visual quality of HEVC compression artifacts to five visual quality levels. The original VMAF classification was changed to reflect High Definition (HD) resolution images. We trained a DenseNet network to classify compressed images into five visual categories using the dataset generated by the modified VMAF. DenseNet was chosen for its ability to process HD images. Our evaluations have shown that our no-reference metric achieves an impressive average accuracy of 94.13% in classifying the visual quality of compressed images.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".