Simplification of polyline bundles of graphs and trees
Bibliographic record
Abstract
Polyline simplification is a well-studied optimization problem, in which a given polyline shall be replaced by a polyline with fewer vertices which still represents the shape of the original polyline faithfully. In this paper, we propose and study a generalization of the polyline simplification problem. Instead of a single polyline, we are given a set of $\ell$ polylines possibly sharing some line segments and vertices. We call such a set a polyline bundle. The task is to simplify each polyline $L$ of a given polyline bundle by keeping a subset of its vertices such that (i) the Hausdorff or Fréchet distance between $L$ and its simplified counterpart does not exceed a given distance threshold $\delta$, (ii) a shared vertex is either kept or discarded in all polylines of the polyline bundle (we refer to this requirement as consistency) and (iii) the number of kept vertices in the polyline bundle is minimized. To justify this definition, we argue that consistency is crucial to get meaningful and aesthetically pleasing outputs. Regarding the computational complexity of polyline bundle simplification, we prove that this problem is NP-hard to approximate within a factor of $n^{1/3−\varepsilon}$ for any $\varepsilon > 0$, where $n$ is the number of vertices in the polyline bundle. This inapproximability even applies to planar inputs and also to instances with only $\ell=2$ polylines. However, we identify the sensitivity of the solution to the choice of the distance threshold $\delta$ as a reason for this strong inapproximability. In particular, we prove that if we employ the Fréchet distance and allow $\delta$ to be exceeded by a factor of $2$ in the solution, then we can find a simplified polyline bundle with no more than $O(\log(\ell + n)) \cdot \mathrm{OPT}$ vertices in polytime, providing us with an efficient bi-criteria approximation. In addition, we show that the polyline simplification problem is solvable in polytime in case the polylines form a rooted tree. We further present a greedy heuristic that decomposes general bundles into tree bundles, which then can be simplified individually and optimally. In our experimental study, we compare the performance of the bi-criteria approximation algorithm and the tree bundle decomposition algorithm on public transit networks and movement trajectories. We show that in case the polylines form grid-like structures, the bi-criteria approximation algorithm outputs smaller simplifications, but the tree bundle decomposition algorithm scales better and produces superior results on polyline bundles derived from paths in embedded road networks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".