Utilizing NLP Sentiment Analysis Approach to Categorize Amazon Reviews against an Extended Testing Set
Bibliographic record
Abstract
Sentiment analysis, also known as opinion mining, is a pivotal aspect of natural language processing (NLP). This method entails discerning the polarity of textual information and determining whether it conveys positive or negative sentiments. In one of the domains, e-commerce, sentiment analysis assumes paramount significance. It offers businesses a nuanced understanding of their brand and product sentiment as reflected in customer reviews, facilitating market comprehension and strategic decision-making. This study primarily focused on analyzing the Amazon food reviews dataset, augmenting the original dataset with newly generated data, and subsequently conducting data preprocessing tasks, encompassing text cleansing, removing stop words, lemmatization, and stemming. Subsequently, machine learning models were constructed, trained, and evaluated using NLP feature extraction techniques to address the sentiment analysis challenge and investigate the impact of increased data volume on model performance. Among the diverse methodologies employed for extracting features from textual data samples, this research integrated term frequency-inverse document frequency (TF-IDF), Word to Vector (W2V), and Bag of Words (BoW) techniques in the feature extraction phase. Furthermore, three distinct machine learning models, namely Logistic Regression, Decision Tree, and Random Forest, were designed, implemented, and assessed. The models' performance was scrutinized following hyperparameter optimization to determine the most effective approach. The outcomes revealed that the performance of the models was consistent, yielding accuracy rates ranging from 85% to 89% on the testing dataset. Nevertheless, the Logistic Regression model, employing BoW features, demonstrated superior performance compared to the other models. Following optimization of the logistic regression model, a remarkable accuracy of 89% was attained on the testing dataset by operating the BoW extracted features.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.005 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".