Lossless Compression of Grayscale and Colour Images Using Multidimensional CSE
Bibliographic record
Abstract
Originally, compression by substring enumeration (CSE) is a lossless compression technique that is intended for strings of bits. As such, the original version is one-dimensional. An extension of CSE for strings drawn from a larger alphabet has later been introduced. Also, CSE has recently been extended to two-dimensional (2D) data. As such, 2D CSE can be used directly to compress images. Unfortunately, CSE generally does not perform on data drawn from large alphabets as well as on binary data. This means that, although we can expect 2D CSE to perform well on bilevel images, we must expect a loss of performance on grayscale and colour images, where the alphabet sizes may be 28and 224, respectively, as in common image formats. As a workaround for this difficulty, we propose to handle grayscale and colour images by remaining in the realm of binary data but by extending CSE to higher dimensions. Grayscale images may have the levels of gray of their pixels decomposed into bit planes and, then, get compressed using a 3D CSE. Colour images may have their three colour channels treated as yet another dimension and, then, get compressed using a 4D CSE. Actual empirical measurements are deferred to another paper as we do not have a working implementation of multidimensional CSE yet.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".