Detecting Deepfakes using Temporal Consistency of Facial Expression Transitions
Bibliographic record
Abstract
Deepfake generation techniques have developed at a rapid rate, making it possible to generate highly realistic yet misleading videos with potentially far-reaching implications for privacy, security, and public confidence. This paper presents a study on the detection of deepfakes using the temporal consistency of facial expression transitions. Our method captures and integrates significant spatial and temporal information, facial edges, and dense optical flow with an Xception-based CNN and a bidirectional LSTM (BiLSTM) with an attention mechanism. We evaluated the approach on a multi-expression dataset obtained from DeeperForensics-1.0, comparing performance systematically across a range of expressions from Angry to Neutral. The experiments demonstrate a detection rate of up to $98.38 \%$ on the combined multi-expressions and point to the unique challenge of less expressive emotions. The findings affirm that face expression continuity examination plays an important part in enhancing the robustness of deepfake detection, achieving a scalable and adaptive approach to verifying the integrity of real-world media.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".