Retail Customer and Market Proclivity Assessment using Historical data and Social Media Analytics
Bibliographic record
Abstract
Predictive analytics is a field which enables to predict the various aspects of a business. It offers vast prospects in today's business transformation by delivering an automated decision-making process. We focus our research on gaining insights into the retail market operations. Therefore, we introduce an integrated proclivity assessment model to examine the influence of consumer-market relationships. First, we perform a qualitative study to investigate the potential sales of the grocery market by utilizing the consumer's transactional history data released by Instacart. Our ensemble model for predicting purchase probabilities of the products performs better than the currently used baseline algorithms by achieving 88.84% accuracy. Based on the developed lexicon and rule-based sentiment analysis tool, our second proposed solution proficiently interpret the user propensity (77.76%) towards grocery products by scrutinizing consumer tweets based on the user location. Finally, the third model which proposes an extended framework to detect the category of purchase intentions reflected by the consumers in their online reviews shows the highest F1 score of 94.17%. If the information about the product experience is present explicitly in the review data, our customized technique can accurately segregate the different purchase intension labels (positive, negative, and unknown).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".