Bibliographic record
Abstract
Content-defined chunking (CDC) based data deduplication is a complex process, leading to the use of rule-of-thumb approaches and standardized parameter values. However, our work challenges these standard approaches which can lead to worse deduplication ratios, and reemphasizes that parameters need to be optimized for each dataset. We expose new pitfalls, analyze the behaviour of the underlying deduplication process, and provide solutions to aid future deduplication work.Deduplication research often solely reports the expected chunk length of their system without providing the low-level parameters. Our results show that expected chunk length is inadequate for properly describing deduplication, making empirical reproducibility challenging. In fact, different parameter sets with the same expected chunk length can yield different deduplication ratios. We further show that this discrepancy can be explained by chunk length variance.We find that because expected average and standard deviation of chunk length do not account for file boundaries and fingerprint value noise, they can significantly differ from observed values. We show that these phenomena can also cause the same parameter set to give different chunk length distributions on different datasets, explaining why parameters must be tuned to each dataset. However, we find on our datasets that the maximum chunk length parameter does not need to be tuned, and that a rule-of-thumb value of 2<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">16</sup> is a reasonable selection. We provide our datasets in full to promote reproducibility and further work.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".