A Multi-task Spanish Speech Emotion Recognition System for Intelligent Human-Computer Interaction
The core document of the system: architecture, training procedure, and the multi-task results reported above.
A frozen Wav2Vec2 XLSR encoder trained across six Spanish-language corpora to jointly predict emotion, speaker profile, and regional accent — built as the perception layer for empathic human-computer interaction.
Reported without data augmentation or resampling.
The Wav2Vec2 XLSR encoder remains frozen; three lightweight task heads are trained on top of it to predict emotion, speaker profile, and regional accent from a single shared representation of the input speech.
Six Spanish corpora are combined to cover a wider range of speaking styles, speakers, and recording conditions than any single dataset provides.
The recognition model is the core contribution of this research; two further modules show how it operates inside a complete spoken system — from listening to responding.
The multi-task model at the center of this research: a single encoder trained to recognize emotion, speaker profile, and regional accent from Spanish speech.
Visit module →
Speech synthesis that renders responses with an emotional tone consistent with the conversation, rather than a flat, neutral delivery.
Visit module →
Automatic speech recognition, emotion recognition, an LLM guided by a response protocol, and emotional text-to-speech, connected in a single conversational flow.
Visit module →Three publications document the system, from its architecture to its role in a benchmark study and a conference presentation.
The core document of the system: architecture, training procedure, and the multi-task results reported above.
A benchmark of pre-trained models for Spanish speech emotion recognition, forming the empirical basis for the encoder used in the thesis.
A conference presentation of the multi-task framework and its results.
Universidad Veracruzana
Faculty of Electronic Instrumentation
INAOE
SECIHTI