A deep generative model framework for creating \nhigh quality synthetic transaction sequences
Bibliographic record
Abstract
Synthetic data are artificially generated data that closely model real-world measurements, \nand can be a valuable substitute for real data in domains where it is costly \nto obtain real data, or privacy concerns exist. Synthetic data has traditionally been \ngenerated using computational simulations, but deep generative models (DGMs) are \nincreasingly used to generate high-quality synthetic data. \nIn this thesis, we create a framework which employs DGMs for generating highquality \nsynthetic transaction sequences. Transaction sequences, such as we may see in \nan online banking platform, or credit card statement, are important type of financial \ndata for gaining insight into financial systems. However, research involving this type \nof data is typically limited to large financial institutions, as privacy concerns often \nprevent academic researchers from accessing this kind of data. Our work represents \na step towards creating shareable synthetic transaction sequence datasets, containing \ndata not connected to any actual humans. \nTo achieve this goal, we begin by developing Banksformer, a DGM based on the \ntransformer architecture, which is able to generate high-quality synthetic transaction \nsequences. Throughout the remainder of the thesis, we develop extensions to Banksformer \nthat further improve the quality of data we generate. Additionally, we perform \nextensively examination of the quality synthetic data produced by our method, both \nwith qualitative visualizations and quantitative metrics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.004 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".