Cost-sensitive PLM-based Approach for Arabic and English Cyberbullying Classification
Résumé
Abstract Cyberbullying is a hurtful phenomenon that spreads widely on social networks and negatively affects the lives of individuals. Detecting this phenomenon is of utmost necessity to make the digital environment safer for youth. This study uses a bilingual classification of cyberbullying on Ara-bic and English datasets. A four-module approach is proposed. It consists of preprocessing the textual data, generating sentence embeddings, performing the classification, and evaluating the results of the models. The approach relies on two strategies based on transfer learning of pre-trained NLP models. The first uses PLMs (ELMo, Universal Sentence Encoder, BERT, distilBERT, and RoBERTa) to generate sentence embeddings, while the second adopts a fine-tuning procedure of BERT-based PLMs for cyberbullying classification. Due to the frequent class imbalance problem in the research literature, this study used cost-sensitive learning algorithms trained to maximize the Recall/F1 score. The aim is to search for the best classification model that most accurately separates the cyber-bullying and non-cyberbullying classes. The models achieve 75-84%.
Citer ce document
Exporter : BibTeX · RIS (Zotero, Mendeley, EndNote)
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueLicence et provenance
Licence : CC BY
Notice moissonnée depuis OpenAlex le 10/09/2026. Le document reste hébergé par sa source.
Voir le document à la source →
Statistiques
Consultations : 3
Téléchargements : 0