Generalised Additive Modelling of Auto Insurance Data with Territory Design: A Rate Regulation Perspective
Bibliographic record
Abstract
Pricing using a Generalised Linear Model is the gold standard in the auto insurance industry and rate regulation. Generalised Additive Model applications in insurance pricing are receiving increasing attention from academic researchers and actuarial pricing professionals. The actuarial practice has constantly shown evidence of significantly different premium rates among the different rating territories. In this work, we build predictive models for claim frequency and severity using the synthetic Usage Based Insurance (UBI) dataset variables. First, we conduct territorial clustering based on each location’s claim counts and amounts by grouping those locations into a smaller set, defined as a cluster for rating purposes. After clustering, we incorporate these clusters into our predictive model to determine the risk relativity for each factor level. Through predictive modelling, we have successfully identified key factors that may be helpful for the rate regulation of UBI. Our work aims to fill the gap between individual-level pricing and rate regulation using the UBI database and provides insights on consistency in using traditional rating variables for UBI pricing. Our main contribution is to outline how GAM can address a more complicated functionality of risk factors and the interactions among them. We also contribute to demonstrating the territory clustering problem in UBI to construct the rating territories for pricing and rate regulation. We find that relativity for high annual mileage driven is almost three times that associated with low annual mileage level, which implies its importance in premium calculation. Overall, we provide insights into how UBI can be regulated through traditional pricing factors, additional factors from UBI datasets and rating territories derived from basic rating units and the driver’s location.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".