Analysis of Scanline and Minimum Entropy Selection Heuristics in Model Synthesis and Wave Function Collapse
Bibliographic record
Abstract
In this paper, we analyze two versions of a texture synthesis algorithm, study the cases in which they fail to produce a successful result and present modifications that could be made to lessen their rates of failure. This algorithm, Model Synthesis, and its variation Wave Function Collapse are designed to take in a small sample input image, or set of image constraints, and produce a larger pseudorandom output image in which every region of the output image is locally similar to an element of the input image. Both versions of the algorithm accomplish this task by considering their output image as a grid of cells with each cell initially in a superposition of all possibilities for itself and resolving cells one by one until all cells have been resolved from their superposition to a fixed value. One of the key differences between the two versions of the algorithm is the order in which cells are selected to be resolved, one simply selects in a scanline order, while the other resolves first those cells that have the minimum entropy, and thus which we can be most certain of their eventual state. For many inputs, the minimum entropy model reaches a state in which its output is not consistent with the input and thus fails, while the scanline model does not. This paper looks at the cases in which this occurs and concludes that this is often caused by the minimum entropy model creating regions of elevated constraints in its solution. Finally, it presents a possible alteration to the algorithm which allows a minimum entropy model to avoid this manner of failure among a subset of test cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.036 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".