A Generic Form for Capturing Unobserved Heterogeneity in Discrete Choice Modelling: Application to Neighborhood Location Choice
Bibliographic record
Abstract
Discrete choice models and their strength to predict individual choices mostly depend on the quality of datasets that have been used for model generation. However, even the most comprehensive and detailed datasets are not able to observe all factors pertinent to someone’s choice. This issue in the choice modelling literature has been addressed as unobserved heterogeneity, which means that individuals across populations are not affected identically by alternative attributes. Furthermore, such variation in preferences across populations and their sources are not always recognized by researchers. \nThere are different methods to capture unobserved heterogeneity proposed in the discrete choice literature among which the random parameters approach, also referred to as mixed logit models, the latent class approach and the agent effect approach are the most well know methods. The main contribution of this study is to extend the formulation of LC-MMNL model to capture the agent effect by including a random term in the utility function of the model. Three types of models, Mixed Multinomial Logit (MMNL), Latent Class Mixed Multinomial Logit (LC-MMNL) and Agent Effect Latent Class Mixed Multinomial Logit (AGLC-MMNL) have been generated and the results compared. Considering agent effect simultaneously with other sources of unobserved heterogeneity in a latent class context demonstrates improvement in terms of model fit as well as cross section validation. It enables us to generate a latent class model with a larger number of classes explaining more heterogeneity across the population of a neighborhood location choice study. The AGLC-MMNL model is able to detect four distinct classes of individuals in Montreal, exhibiting different behaviours while facing neighborhood location choices in the context of a Discrete Choice Experiment. The classes of the model are able to explain different behaviours of individuals based on their income level, whether they are transit or car oriented, and the importance of privacy to them.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".