A Hybrid RNN-LSTM Framework for Predicting Stakeholder Engagement on Social Media Platforms
Bibliographic record
Abstract
Social media has turned into a vital means of communication for organizations to engage with stakeholders, although the prediction of engagement is much more difficult due to the nature of dynamic and sequential interactions between users.Social media is a vital means to fostering relations with stakeholders; yet, accurate engagement prediction is still not straightforward, given its dynamic and sequential user interaction.In this paper, we propose a resource-efficient temporal deep learning architecture that leverages a low-cost Recurrent Neural Network (RNN) and stacked Long Short-Term Memory (LSTM) units for addressing both short-term contextual influences and long-term engagement trends in social media performance.Experiments were conducted on a balanced multi-platform dataset of media activity.Experiments were conducted on a balanced multi-platform dataset of 100k posts, including textual content, engagement metrics and metadata.The data were rigorously preprocessed, including cleaning, tokenization, stop word removal and term frequency inverse document frequency (TF-IDF) vectorizing but keeping the top 1,000 informative features to make it computationally efficient and interpretable.The proposed model employs sequential LSTM layers, with dropout regularization and trained using the Adam optimizer for a small number of epochs to avoid overfitting.Empirical experimentation on the test set held out from the training shows that our model has strong predictive performance achieving 99.60% accuracy, 99.61% precision, 99.60% recall, and an AUC score of 1.0 for binary engagement classification.Although the findings demonstrate that time sequence modeling is effective in engagement prediction, the paper prioritizes efficiency and practical deployability rather than architectural complexity.The proposed system lays a scalable foundation for stakeholder analytics, content optimization and engagement-aware recommendation systems and future works will focus on crossplatform generalization as well as multimodal expansions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".