RKLT: 8 DOF Real-Time Robust Video Tracking Combing Coarse Ransac Features and Accurate Fast Template Registration
Bibliographic record
Abstract
The performance of a tracker can be measured by two often conflicting criteria - robustness and accuracy. Recently researchers have focused on improving robustness, using adaptive appearance models. However updating the appearance model can cause drift and lower the accuracy of motion (state) estimation. These trackers generally compute 2 degree of freedom(DOF) image translation of the object, and are suited for applications such as surveillance. In contrast, we are interested in tracking objects using high DOF motion models - especially 8DOF homograph models that allow tracking of precise state information (projective 8D or calibrated 3Dworld translations and 3D rotations of the tracked object). Such precise state is required for visual motion control of e.g. robot arms, hands and UAV. To this end, we propose a novel tracking algorithm that combines KLT [8], RANSAC [21] and Inverse Compositional tracker [7]. First we sample a large patch into a set of small patches and track each one using frame-to-frame 2D KLT trackers. An 8 DOF homograph describing the large patch motion is then estimated from the current locations of these KLT trackers using RANSAC while also discarding lost trackers as outliers. Finally, using the RANSAC 8DOF motion estimate as the initial guess, we perform a few iterations of an IC registration tracker. This refines the patch motion to sub-pixel accuracy and avoids drift by registering to the original template. We perform three sets of experiments - one is the standard synthetic Lena convergence benchmark and two use real image sequences from recent datasets - to show that our tracker compares favourably with the state-of-the-art.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".