Video Data Extraction and Processing for Investigation of Vehicles’ Impact on the Asphalt Deformation Through the Prism of Computational Algorithms
Bibliographic record
Abstract
There are numerous algorithms and solutions for car or object detection as humanity is aiming towards the smart city solutions. Most solutions are based on counting, speed detection, traffic accidents and vehicle classification. The mentioned solutions are mostly based on high-quality videos, wide angles camera view, vehicles in motion, and are optimized for good visibility conditions intervals. A novelty of the proposed algorithm and solution is more accurate digital data extraction from video file sources generated by security cameras in Bosnia and Herzegovina from M18 roadway, but not limited only to that particular source. From the video file sources, data regarding number of vehicles, speed, traveling direction, and time intervals for the region of interest will be collected. Since finding contours approach is effective only on objects that are mobile, and because the application of this approach on traffic junctions did not yield desired results, a more specific approach of classification using a combination of Histogram of Oriented Gradients (HOG) and Support Vector Machines (Linear SVM) has shown to be more appropriate as the original source data can be used for training where the main benefit is the preservation of local second-order interactions, providing tolerance to local geometric misalignment and ability to work with small data samples. The features of the objects within a frame are extracted first by standardizing the feature variables and then computing the first order gradients of the frame. In the next stage, an encoding that remains robust to small changes while being sensitive to local frame content is produced. Finally, the HOG descriptors are generated and normalized again. In this way the channel histogram and spatial vector becomes the feature vector for the Linear SVM classifier. With the following parameters and setup system accuracy was around 85 to 95%. In the next phase, after cleaning protocols on collected data parameters, data will be used to research asphalt deformation effects.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".