UTILIZING TEXT SENTIMENT AUTOMATIC BENGALI BULLYING DETECTION THROUGH MACHINE LEARNING
No Thumbnail Available
Date
2022-08-30
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
The number of Bengali language users on social media is rapidly increasing, and the use
of code-mixing and transliteration. Multiple languages mix using code-mixing.
Transliterated Bengali allows writing Bengali with Latin words. Manually reviewing and
removing bullying content from social media can be time-consuming, which is
undesirable in today's technologically automated world. The machine learning approach
can be convenient in keeping the system updated with new types of abuser approaches.
Machine learning algorithms such as a k-nearest neighbor, Naive Bayes, Support Vector
Machine, Random Forest, Logistic Regression, and AdaBoost classification were applied
to distinguish transliterated Bengali cyberbullying content. The dataset contained not only
transliterated Bengali text but also Bengali and Code-Mixed Bengali text. Baseline
features from the social media dataset were used with sentiment and personality features.
The features were the number of likes and dislikes, profane words, positive words,
negative words, related to posts, and sentiment. TF-IDF and n-grams technique was used
for feature extraction. The performances of the algorithms were evaluated utilizing
precision, recall, accuracy, F1 score, AUC, and ROC curve. Among sentiment analysis.
The result shows that the linear SVM algorithm achieved the highest accuracy of 64.224%
to classify sentiment for test data. In contrast, the SVM and RF outperformed other
classifiers in detecting bullying comments. For testing data, the SVM and RF achieved
94.83% and 94.40% accuracy, respectively. This work will have an impact on reducing
transliterated and code-mixed Bengali cyberbullying.