Application of a Machine Learning Algorithm to Assess and Minimize Credit Risks
Bibliographic record
Abstract
The banking system, as the most important sector of the economy of every country, often encounters a number of risks. Financial institutions of that system operate in an unstable environment, and without having complete information about that environment, they may suffer significant losses. The main source of such losses is considered to be credit risks, and for the management of these, various mathematical models are being developed which will allow banks to make decisions on granting a loan. Lately, for this purpose, machine learning (ML) classification algorithms have often been used for credit risk modeling. In this research work, using the ideas of well-known ML algorithms, a new algorithm for solving the binary classification problem was developed. By means of the algorithm created, based on real data, a classification model has been developed. Qualitative indicators of that model, such as ROC AUC, PR AUC, precision, recall, and F1 score, were evaluated. By modifying the resulting probabilities into a range of 300–850 score points, a scoring model has been developed, the usage of which can mitigate credit risk and protect financial organizations from major losses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".