Deep Metric Learning for Near-Duplicate Video Retrieval Leveraging Efficient Semantic Feature Extraction
Bibliographic record
Abstract
Video sharing platforms like YouTube, TikTok and Instagram have gained popularity in the online space. Daily several videos are uploaded, which calls for an efficient video retrieval system that could identify near-duplicate videos that offers several advantages in content management, copyright protection, and multimedia retrieval. This will facilitate efficient content management by removal of redundant videos from large repositories to streamline storage resources and improve accessibility of multimedia collections. Additionally, this can help copyright protection and intellectual property allowing right holders to identify unauthorized copies of their original work. Moreover, in applications such as multimedia retrieval and recommendation systems, removal of near-duplicate videos can enhance user experience by providing relevant search results. AI provides a promising solution to this problem. We have proposed an effective system built on deep metric learning that solves the near duplicate video retrieval. This proposed model uses the pre-trained VGG-16 network that contains convolutional and fully connected layers to find video features. These video representations are fed to the deep metric learning framework in the form of triplets which are trained to calculate the accurate distance between similar or near-duplicate videos. For the training of the framework, VCDB dataset was used whereas for the evaluation of the model CC_WEB_VIDEO and TRECVID BBC Rushes 2007 datasets were used. Experiments have shown that mean average precision of 0.985% for the CC_WEB_VIDEO dataset is achieved thus outperforming the state-of-the-art models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".