A comprehensive approach for evaluation of \nstereo correspondence solutions \nin outdoor augmented reality
Bibliographic record
Abstract
For many years, researchers have made great contributions in the fields of augmented \nreality (AR) and stereo vision. One of the most studied aspects of stereo vision \nsince the 1980s has been Stereo Correspondence, which is the problem of finding the \ncorresponding pixels in stereo images, and therefore, building a disparity map. As \na result, many methods have been proposed and implemented to properly address \nthis problem. Due to the emergence of different techniques to solve the problem \nof stereo correspondence, having an evaluation scheme to assess these solutions is \nessential. Over the past few years, different evaluation schemes have been proposed \nby researchers in the field to provide a testbed for assessment of the solutions based \non specific criteria. Middlebury Stereo and Kitti Stereo benchmarks are two of the \nmost popular and widely used evaluation systems through which a solution can be \nevaluated and compared to others. However, both of these models take a general \napproach towards evaluating the methods, that is, they have not been designed with \nan eye to the particular target application. In our proposed approach, steps are taken \ntowards an evaluation design based on the potential applications of stereo methods, \nwhich enables us to better define the criteria for efficiency, that is, the processing \ntime, and the required accuracy of the disparity results. Since AR has attracted \nmore attention in the past few years, the evaluation scheme proposed in this research \nis designed based on outdoor AR applications which can take advantage of stereo \nvision techniques to obtain a depth map of the surrounding environment. This map \ncan then be used to integrate virtual objects in the scene that respect the occlusion \neffects that are expected to occur based on the depth of the real objects.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.005 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".