Image Compression for Communication: Topic Modelling of Scopus Abstracts Using BERTopic
Bibliographic record
Abstract
In the era of increasing digital data generation and utilization, efficient data compression is increasingly crucial, particularly in the realm of image and video communication. Visual content continues to be extremely important in digital platforms such as social media, streaming services, video conferencing, and medical images. Management of the substantial daily data flow poses significant challenges for storage, transmission, and real-time processing. Due to their inherently large size, image and video files necessitate effective compression techniques to optimize bandwidth usage, reduce storage requirements, and facilitate the seamless delivery of content across diverse communication networks. In this study, we utilized AI based topic modeling using BERTopic for a review of $\mathbf{5 0, 2 6 1}$ Scopus abstracts published from 1972 to mid-2024 to investigate dominant topics of research done on image data compression. The variability in the results of the BERTopic model is observed in multiple runs on same model parameters and is taken care of by taking 10 runs and keeping stable topics that are found in multiple runs. The overall topics of all the experiments are finalized into 10 major topics, namely, image compression, watermark encryption and security, MIMO (multiple input multiple output) channels system, video compression, deep neural networks for image compression, saliency detection, medical PCS system, satellite broadcasting and holograms in communication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".