Two Essays on Analyst Information Processing
Bibliographic record
Abstract
This thesis consists of two essays focusing on how sell-side analysts provide value to the market through their information processing.The first essay studies the impact of private interaction with management, one of the most crucial information sources for sell-side analysts, on their performance.I identify private interaction between sell-side analysts and firm management based on 373,869 analyst reports for 2,958 U.S. firms by 2,238 analysts who work with eight large banks during 1997-2019.I find that earnings forecast accuracy increases after an analyst privately interacts with management, especially when the information asymmetry and forecast difficulty of the underlying firms are higher.Private interaction with management helps analysts reduce bias, produce more detailed and soft information such as operational and strategic information, and achieve better career outcomes.My findings provide direct evidence that private interaction with management benefits analyst performance.The second essay investigates the impact of subjectivity, one of the most common attributes in textual information, on the informativeness of analyst reports.We use machine learning techniques to classify statements in 421,583 analyst reports into objective facts and subjective opinions.We find that market reaction to analyst narratives increases with analysts' subjectivity in their research reports.The effects are more pronounced for firms with lower financial report quality and for analysts' subjective assessment on risk and growth, but less pronounced when macro uncertainty is higher.We also find that higher prevalence of analysts' subjectivity is associated with better future firm earnings growth.Additional analyses indicate that analyst report subjectivity captures analysts' additional effort allocation and private information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.004 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".