Emozionalmente: A Crowdsourced Corpus of Simulated Emotional Speech in Italian
Fabio Catania, Jordan W. Wilke, Franca Garzotto · 2025 · 5 citationsRead the paper
Speech emotion recognition (SER) relies on speech corpora collecting emotional voices for analysis. Emotions may vary by culture, language, and context—whether simulated, induced, or naturalistic—and resources featuring simulated emotional speech in Italian are scarce. To address this gap, we launched a crowdsourcing campaign and obtained Emozionalmente, an acted corpus of 6902 samples produced by 431 non-professional Italian speakers verbalizing 18 sentences simulating the Big Six emotions and neutrality. A subjective validation involving 829 individuals achieved a recognition accuracy of 66%, demonstrating the corpus' utility and representativeness. We also explored SER using the pretrained deep learning model wav2vec 2.0, fine-tuned to our corpus, which attained an accuracy of 82.45%. This paper details the crowdsourcing methodology, corpus analysis, and both human and computational validation studies, highlighting the corpus' broad applications in SER. Emozionalmente is publicly available, supporting further research in linguistics, speech processing, and affective computing within the Italian context.
1 idea Seedlabs derived from this research
A specialized B2B testing tool for developers of Italian voice assistants to validate whether their synthetic voices sound emotionally authentic to native speakers.
AI score 84/100