EmotionNet: A Novel Hybrid Deep Learning Model for Arabic Speech Emotion Recognition
Résumé
This study presents EmotionNet, a novel hybrid deep learning model designed for Arabic speech emotion recognition. EmotionNet integrates a Variational Auto-Encoder (VAE) for latent representation learning with a lightweight classification branch enhanced by latent-space refinement. Evaluated on the KEDAS dataset, which includes five emotionally acted categories, the model achieved a test accuracy of 93.99% and outperformed conventional classifiers such as SVM, MLP, and Random Forest. The proposed approach employs a compound loss function and KL annealing to jointly optimize reconstruction and classification. Although the results are promising, the acted nature of KEDAS may overstate real-world performance, highlighting the need for evaluation on spontaneous, multimodal datasets, an effort currently underway in an ongoing interdisciplinary project.
Citer ce document
Exporter : BibTeX · RIS (Zotero, Mendeley, EndNote)
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueLicence et provenance
Licence : CC BY
Notice moissonnée depuis OpenAlex le 29/09/2026. Le document reste hébergé par sa source.
Voir le document à la source →
Auteur(s)
Statistiques
Consultations : 1
Téléchargements : 0