Marlin: Taming the big streaming data in large scale video similarity search
Bibliographic record
Abstract
The extreme volume and staggeringly increasing rate inevitably produce unprecedented pressure on any large scale video sharing and hosting systems. Among the efforts to mitigate this pressure, content-based video similarity search is becoming more and more important with the exponential growth of the data size. Though various approaches have been proposed to address this problem, they are mainly focusing on the retrieval accuracy thus bringing video features with high complexity. Due to the complexity of the feature, these systems are based on the assumption that features representing videos have been obtained offline and stored in the database statically. However, the on-call efforts to move the feature extraction and similarity search from offline to online have been ignored in previous work. In this paper, we propose Marlin, a streaming data processing pipeline that efficiently extracts video features and retrieves video similarity information in a large scale video data system. We design a streaming feature extractor to handle the videos streaming into the system and establish the fined-grained resource allocation with a resource-aware data abstraction layer over streaming data to allocate computing resources among the videos with various resource demands. Besides that, we are pipelining the feature extraction and similarity search process with a distributed feature index, which supports real-time query and incremental index update. The experimental and the extensive real-world workload driven simulation results show that the proposed stream processing architecture achieves 25X speedup against the sequential feature extraction algorithm and 23X speedup against the sequential similarity search with a subsecond similarity query latency for a single request.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".