Bibliographic record
Abstract
Intelligent/autonomous vehicles, such as self-driving cars, intelligent robots and Unmanned Aerial Vehicles (UAVs) must seamlessly interact with humans, e.g., their drivers/operators/pilots or people in their vicinity, whether being obstacles to be avoided (e.g., pedestrians) or targets to be followed and interact with (e.g., when filming a performing athlete). Furthermore, intelligent vehicles and robots have been increasingly employed to assist humans in real-world applications (e.g., for, autonomous transportation, warehouse logistics, or infrastructure inspection). To this end, autonomous vehicles should be equipped with advanced vision systems that allow them to understand and interact with humans in their surrounding environment. This lecture overviews humancentric AI methods that can be utilized to facilitate visual interaction between humans and autonomous vehicles (e.g., through gestures captured by RGB cameras), in order to ensure their safe and successful cooperation in real-world scenarios. Such methods should: a) demonstrate increased visual perception accuracy to understand human visual cues, b) be robust to input data variations, in order to successfully handle illumination/background/scale changes that are typically encountered in real-world scenarios, and c) produce timely predictions to ensure safety, which is a critical aspect of autonomous vehicles' applications. Deep learning and neural networks play an important role towards this end, covering the following topics: a) human pose/posture estimation from RGB images, b) human action/activity recognition from RGB images/skeleton data, and c) gesture recognition from RGB images/skeleton data. Finally, embedded execution is extremely important, as it facilitates vehicle autonomy, e.g., in communication-denied environments. Application areas include driver/operator/pilot activity recognition, gesture-based control of autonomous vehicles, or gesture recognition for traffic management. The lecture will offer an overview of all the above plus other related topics and will stress the related algorithmic aspects. Some issues on embedded CNN computation (e.g., through fast convolution algorithms) will be overviewed as well.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.135 | 0.963 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".