MDD-based Domain Adaptation Algorithm for Improving the Applicability of the Artificial Neural Network in Vehicle Insurance Claim Fraud Detection
Bibliographic record
Abstract
Insurance fraud detection is a critical task for insurance companies, as fraudulent claims result in financial losses and increased premiums for honest policyholders. Traditional fraud detection methods rely on rule-based approaches and manual investigation, which are limited in their ability to adapt to evolving fraud patterns. In this study, we propose a novel approach using an artificial neural network (ANN) combined with Margin Disparity Discrepancy (MDD)-based domain adaptation to improve the generalization ability of fraud detection models across different datasets. We first preprocess the data by applying K-Means clustering to segment source and target domains based on distribution differences. We then compare multiple machine learning models, including decision trees, random forests, k-nearest neighbors, and gradient-boosted decision trees, finding that ANN achieves the best performance. To further enhance generalizability, we introduce MDD-based domain adaptation, aligning feature distributions between the source and target domains. Experimental results demonstrate that the adapted ANN significantly improves fraud detection accuracy, achieving a higher F1-score and recall while reducing the false negative rate. These findings highlight the effectiveness of domain adaptation in addressing distributional shifts in fraud detection, making the proposed model a promising solution for real-world insurance fraud detection systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".