The complexity of financial wellness: examining survey patterns via kernel metric learning and clustering of mixed-type data
Bibliographic record
Abstract
Recent market events and inflation have significantly affected the financial stress facing many individuals, but understanding the main stressors is paramount to supporting them in making better long-term financial decisions. Financial advisors must understand the types of stress their clients face to provide tailored advice. While recent high inflation rates may underpin the cause of their clients’ stress, we ask: what are the major sources of stress that affect an individual’s financial wellness? In this study, we analyze the responses of 1874 individuals to 68 mixed-type questions from 2022 using distance-based clustering that is widely used in finance to group data into similar groups. Distance-based clustering is widely used in finance to group data into similar groups, which requires a predefined distance measurement between data points based on their (dis)similarity. We use a mixed-type metric that utilizes a variable-specific kernel functions with cross-validated bandwidths to optimally balance variables important for similarity, and smooth out variables irrelevant to the difference between data points. Applying the metric to the high-dimensional survey, we found two clusters of respondents: (1) the ‘steady savers’, who represent approximately one third of survey respondents and expressed stronger financial well-being with respect to day-to-day financial obligations and future outlooks, and (2) the ‘financial strivers’ who currently find themselves in more financially stressful situations. This segmentation provides financial advisors with useful results to allocate products, services, or advice tailored to support each group’s unique financial wellness needs. By leveraging this methodology, we strive to advance the realm of personalized financial advising and the landscape of robo-advising. Enhanced precision and tailored strategies allow this work to elevate the quality of investment recommendations, contributing to the future of automated financial guidance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".