Cluster analysis of share price: How firm characteristics relate to accounting metrics
Bibliographic record
Abstract
The purpose of this paper is to improve our understanding of the relationship between share price and accounting information. Much of the literature utilizes the earnings number to reflect firm value. However, the revenue number seems more relevant for high tech firms (Xu, Cai, & Leung, 2007), and cash flow figures are more informative for internet companies (Romanova, Helms, & Takeda, 2012). We build on this notion that share price may map out to different accounting numbers for different firms. We collect 629 accounting metrics for 3,365 firms in the U.S. and estimate their correlation with the firms’ share price. We analyze these correlations and find that many firms exhibit a low correlation between share price and earnings. Other accounting numbers are important for these firms, including book value of net assets, retained earnings, stock options, gain or loss items, special or non recurring items, and dividend rates. We are curious to learn what causes firms to anchor onto different metrics, therefore perform a cluster analysis to group similar firms together along three key accounting metrics. We examine the composition of each cluster and find that capital structure, dividend patterns, the persistence of operations, age, and industry can influence which accounting number is correlated with firm value. We encourage other researchers to continue this exploration as there are many interesting questions to answer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.032 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".