{# Audit 04/10/2026 : « autre » n'est pas un code de langue ; SPHAERO n'est pas l'éditeur des documents qu'elle héberge ou référence. #} {# citation_pdf_url doit mener à un PDF : un lien vers une page DOI est pénalisé par Google Scholar (avant : tout lien externe). #}
Accès ouvert · CC BY

Advancing sentiment analysis for low-resourced african languages using pre-trained language models

Article scientifique 2025 Anglais

Résumé

While sentiment analysis systems excel in high-resource languages, most African languages facing limited resources, remain under-represented. This gap leaves a significant portion of the world's population without access to technologies in their native languages. However, multilingual pre-trained language models (PLM) offer a promising approach for sentiment analysis in low-resource languages. Although the absence of large data in African languages poses a challenge for developing PLMs, fine-tuning and task adaptation of existing multilingual PLMs is an alternative solution. This paper explores the use of multilingual PLMs for sentiment analysis in five Southern African languages: Sepedi, Sesotho, Setswana, isiXhosa, and isiZulu. We leverage existing PLMs and fine-tune them for this specific task, avoiding training the models from scratch. Our work expands on the SAfriSenti corpus, a Twitter sentiment dataset for these languages. We employ various annotation techniques to create a labelled dataset and perform benchmark experiments utilising various multilingual PLMs. Our findings demonstrate the effectiveness of multilingual PLM, particularly for closely-related languages (Sotho-Tswana), where the ensemble PLMs method achieved an average weighted F1 score above 63%. In particular, Nguni closely-related languages achieved an even higher average weighted F1 score, exceeding 77%, highlighting the potential of PLMs for sentiment analysis in South African languages.

Citer ce document

Mabokela, K. R., Primus, M., & Çelik, T. (2025). Advancing sentiment analysis for low-resourced african languages using pre-trained language models. PLoS ONE. https://doi.org/10.1371/journal.pone.0325102

Exporter : BibTeX · RIS (Zotero, Mendeley, EndNote)

Accès au document

Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter

Voir l'article sur le site de la revue

Licence et provenance

Licence : CC BY

Notice moissonnée depuis OpenAlex le 28/09/2026. Le document reste hébergé par sa source.
Voir le document à la source →

Statistiques

Consultations : 1

Téléchargements : 0