An Ensemble Approach for Cyber Bullying: Text Messages and Images
Bibliographic record
Abstract
Text mining (TM) is a domain used to find valuable patterns from various text documents.Cyberbullying is the term used to abuse a person online or offline platform.Nowadays, cyberbullying has become more dangerous to people who are using social networking sites (SNS).Cyberbullying is of many types, such as text messaging, morphed images, videos, Etc.It is a challenging task to prevent this type of abuse of the person in online SNS.Finding accurate text mining patterns gives better results in detecting cyberbullying on any platform.Cyberbullying is developed with the online SNS to send defamatory statements or orally bully other persons, or by using the online forum to abuse in front of SNS users.Deep Learning (DL) is one of the significant domains used to extract and learn the quality features dynamically from the low-level text inclusions.In this scenario, Convolution neural network (CNN) are DL models used to train text data, images, and videos.CNN is a compelling approach to preparing these data types and achieving better text classification.This paper describes the Ensemble model with the integration of Term Frequency (TF)-Inverse document frequency (IDF) and Deep Neural Network (DNN) with advanced feature-extracting techniques to classify the bullying text, images, and videos.Feature extraction technique extracts the features of cyber-bullying patterns from the text and images.A limited number of datasets are used to classify the data.The proposed approach also focused on reducing the training time and memory usage, which helps the classification improvement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".