A Hybrid Attention-Driven Recurrent Neural Network Model for Sentiment Classification of Social Media Texts
Notice bibliographique
Résumé
With the rapid expansion of user-generated content on social media platforms like Twitter, Facebook, and Reddit, accurately identifying sentiment from textual data has become an essential yet challenging task due to the informal, noisy, and contextually diverse nature of these platforms. To address this, we propose a Hybrid Attention-Driven Recurrent Neural Network (HA-RNN) model that effectively combines Bidirectional Gated Recurrent Units (Bi-GRU) with a sophisticated attention mechanism for sentiment classification. The model utilizes pre-trained GloVe embeddings (300 dimensions) to capture rich semantic features from raw text, enhancing the initial representation of social media data. The Bi-GRU layers are employed to model sequential dependencies Safeyah Tawil Department of Computer Science and Engineering, Faculty of Information Technology, Zarqa University, Zarqa, Jordan. University of Business and Technology, Jeddah, Saudi Arabia stawil@zu.edu.jo I. INTRODUCTION Social media platforms have become primary channels for individuals to express opinions, emotions, and sentiments on a wide range of topics, including politics, products, services, and global events. The explosive growth of platforms such as Twitter, Facebook, and Instagram has led to an overwhelming amount of unstructured textual data that offers valuable insights into public sentiment [1]. Analyzing this vast content can support businesses, governments, and researchers in understanding user perceptions, improving services, and detecting social trends. in both forward and backward directions, ensuring a comprehensive understanding of context within a sentence. The integrated attention layer enables the model to dynamically focus on sentiment-bearing words, thereby improving classification accuracy and interpretability. We evaluated the proposed model on two widely recognized datasets: the Twitter US Airline Sentiment Dataset and the Sentiment140 Dataset. The HA-RNN achieved an accuracy of 90.8% on the Twitter US Airline dataset and 88.5% on Sentiment140, outperforming traditional models such as CNN (84.3% accuracy), LSTM (86.7%), and Bi-GRU without attention (87.1%). Furthermore, the attention mechanism provided insightful visualization, highlighting the critical words influencing sentiment predictions. The model demonstrated a balanced performance with high precision, recall, and F1-scores, validating its robustness across different sentiment classes. Overall, the HA- RNN model presents an effective and interpretable solution for sentiment analysis on noisy and diverse social media texts, supporting applications in social monitoring, brand analysis, and opinion mining.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,003 | 0,004 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».