NFDI4DS | UHH-SEMS - Publication Details

converting anyone s emotion towards speaker independent emotional voice conversion

FOS: Computer and information sciences Sound (cs.SD) Computer Science - Computation and Language Computer Science - Artificial Intelligence 02 engineering and technology Computer Science - Sound Artificial Intelligence (cs.AI) Audio and Speech Processing (eess.AS) FOS: Electrical engineering, electronic engineering, information engineering 0202 electrical engineering, electronic engineering, information engineering Computation and Language (cs.CL) Electrical Engineering and Systems Science - Audio and Speech Processing

DOI: 10.48550/arxiv.2005.07025 Publication Date: 2020-10-25

Abstract Supplemental Material References Cited by

AUTHORS (4)

Berrak Sisman

Kun Zhou

Mingyang Zhang

Haizhou Li

ABSTRACT

Accepted by Interspeech 2020<br/>Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried out under the assumption that emotion is speaker-dependent. We consider that there is a common code between speakers for emotional expression in a spoken language, therefore, a speaker-independent mapping between emotional states is possible. In this paper, we propose a speaker-independent emotional voice conversion framework, that can convert anyone's emotion without the need for parallel data. We propose a VAW-GAN based encoder-decoder structure to learn the spectrum and prosody mapping. We perform prosody conversion by using continuous wavelet transform (CWT) to model the temporal dependencies. We also investigate the use of F0 as an additional input to the decoder to improve emotion conversion performance. Experiments show that the proposed speaker-independent framework achieves competitive results for both seen and unseen speakers.<br/>

SUPPLEMENTAL MATERIAL

Coming soon ....

REFERENCES ()

CITATIONS ()

EXTERNAL LINKS

OPENAIRE - Products

PlumX Metrics

converting anyone s emotion towards speaker independent emotional voice conversion

RECOMMENDATIONS

FAIR ASSESSMENT

Coming soon ....

JUPYTER LAB

Coming soon ....