Application of a Machine Learning Algorithm to Assess and Minimize Credit Risks
Bibliographic record
Abstract
The banking system, as the most important sector of the economy of every country, often encounters a number of risks. Financial institutions of that system operate in an unstable environment, and without having complete information about that environment, they may suffer significant losses. The main source of such losses is considered to be credit risks, and for the management of these, various mathematical models are being developed which will allow banks to make decisions on granting a loan. Lately, for this purpose, machine learning (ML) classification algorithms have often been used for credit risk modeling. In this research work, using the ideas of well-known ML algorithms, a new algorithm for solving the binary classification problem was developed. By means of the algorithm created, based on real data, a classification model has been developed. Qualitative indicators of that model, such as ROC AUC, PR AUC, precision, recall, and F1 score, were evaluated. By modifying the resulting probabilities into a range of 300–850 score points, a scoring model has been developed, the usage of which can mitigate credit risk and protect financial organizations from major losses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".