News vs. Social Media: Sentiment Impact on Stock Performance of Big Tech Companies
Bibliographic record
Abstract
With the growing prominence of large technology firms and the shift in news dissemination driven by social media, scholars have increasingly examined how public discourse about these companies shapes financial markets. Focusing on Apple, Amazon, and Microsoft during the transitional period of January 2015–January 2020, this study evaluates attention and sentiment across traditional news media, social media, and web search in relation to stock market outcomes. We use relatively fine-grained weekly data to link media attention and sentiment to stock returns, volatility, and trading volume. To compare media sentiment across sources, we apply FinBERT-based sentiment analysis, drawing on advances in domain-specific language modeling tailored to financial texts. Results show that social media sentiment (Twitter), exerts a consistently positive and significant influence, while the effects of traditional news media (New York Times) and web search activity (Google Trends) are more irregular. The impact also varies across firms: Twitter sentiment is strongly related to trading volume and volatility for Amazon and Microsoft, but appears less influential for Apple, whose large trading base may dilute the effect. These findings offer a historical baseline for media–finance interactions and highlight how text analysis illuminates the pre-COVID era of big technology firms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".