Detecting and Understanding Sentiment Trends and Emotion Patterns of Twitter Users—A Study on the Demise of a Bollywood Celebrity
Bibliographic record
Abstract
Detecting societal sentiment trends and emotion patterns is of great interest. Due to the time-varying nature of these patterns and trends this detection can be a challenging task. In this paper, the emotion patterns and trends are detected among social media users in a certain case and it is noted that the detection of the trends and patterns is especially difficult in this medium because of the use of informal language. In particular, the role of social networks in the expression of emotions relating to the death of a well-known and loved Bollywood actor Sushant Singh Rajput (SSR) by their fans is explored. The data for the analysis of the emotional state and the sentiment levels of the fans has been acquired from Twitter posts. Different existing sentiment analysis algorithms were compared for the study and chosen for identifying the sentiment trend over a specific timeline of events. The same Twitter posts were also analyzed for emotional content by extracting linguistic features using the psycholinguistic package, Linguistic Inquiry and the Word Count package (LIWC), relating to emotions. Additionally, viral hashtags extracted from the Twitter posts have been segmented and analyzed in order to identify new viral hashtags expressed by the posts over time. The associations between the old and new viral hashtags and between sentiment trends and emotional shifts among the fan base of SSR have been determined and presented graphically.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".