Risco de insolvência e sentimento textual bancário: uma análise dos bancos de capital aberto no Brasil
Bibliographic record
Abstract
This study aimed to analyze whether the textual sentiment explains the greater risk of insolvency of publicly held banks in Brazil, with a sample composed of 17 companies and 450 observations referring to the period from the fourth quarter of 2012, until the fourth quarter of 2019. The empirical strategy adopted is divided in three parts: the first consists of using the unsupervised algorithm k-means to classify banks according to their risk of insolvency, a new measure of bankruptcy probability was elaborated during this process. At this stage, it was observed that 66 observations were classified as high risk, and 384 as low risk, thus being a more rigorous metric than the Z-score, regarding the classification of banks with high risk of insolvency. Then, supervised machine learning methods naive bayes and random forest and the logit model were used to identify which of these statistical techniques is more robust for the prediction of the variable constructed in the previous step. From the confusion matrix and the accuracy criterion it was possible to identify that the logistic model presented the greatest predictive power. Finally, a third step was taken to assess whether textual sentiment, the real percentage change in Gross Domestic Product (GDP), capitalization, profitability, liquidity, and the size of these firms explain the risk of bank insolvency. The results show that banks with a higher probability of bankruptcy have a more optimistic textual feeling.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".