3D Convolutional Spiking Neural Network for Human Action Recognition Using Modulating STDP With Global Error Feedback
Bibliographic record
Abstract
Video action recognition using 3D Convolutional Neural Networks (CNN) become an increasingly popular strategy in past years with the evolution of machine learning and computer vision. However, the higher memory and computation capacity requirement of these networks leads to the use of low-power memory-saving neural networks to perform video action recognition tasks efficiently. Spike-based information processing and computation of bio-inspired spiking convolutional neural networks perform an essential role when comes to energy efficient memory saving computation for video classification and action recognition tasks which allow on-chip real-time processing. This paper proposes a novel 3D Convolutional Spiking Neural Network (CSNN) architecture with modulating STDP supervised learning via global error feedback for human action recognition in video data. The proposed model includes two 3D convolutional layers, followed by two spiking neuron layers, modeled using Leaky Integrate and Fire (LIF) neurons for feature extraction from video data. Using the modulating STDP learning rule with global error feedback, this model can successfully recognize human actions from video data allowing online parallel computations. The proposed network experimented on two datasets: one 3D image dataset - synthesized 3D MNIST and one video dataset - UCF 101 human action recognition dataset and achieved 71.6% and 63.7% recognition accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".