Beyond Actions: Discriminative Models for Contextual Group Activities
Bibliographic record
Abstract
Human action recognition from realistic videos is a challenging problem in computer vision.Several intrinsic properties such as intra-class variations, background clutter and partial occlusion make it difficult to recognize individual person actions reliably.In this dissertation, we go beyond recognizing individual person actions and focus on group activities instead.This motivates from the observation that human actions are rarely performed in isolation, the contextual information of what other people nearby are doing provides useful cues for understanding the high-level activities.We propose a discriminative model for recognizing group activities.Our model jointly captures the group activity, the individual person actions, and the interactions among them.Two new types of contextual information, group-person interaction and person-person interaction, are explored in a latent variable framework.In particular, we propose two different approaches to model the person-person interaction.One approach is to explore the structures of person-person interaction.Different from most of the previous latent structured models which assume a pre-defined structure for the hidden layer, e.g. a tree structure, we treat the structure of the hidden layer as a latent variable and implicitly infer it during learning and inference.The other approach explores the person-person interaction in feature level.We introduce a new feature representation called the action context (AC) descriptor.The AC descriptor encodes information about not only the action of an individual person in the video, but also the behaviour of other people nearby.Our experimental results demonstrate the benefit of using contextual information for disambiguating group activities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".