Skip to main content

SIG-VC: A SPEAKER INFORMATION GUIDED ZERO-SHOT VOICE CONVERSION SYSTEM FOR BOTH HUMAN BEINGS AND MACHINES

Publication ,  Conference
Zhang, H; Cai, Z; Qin, X; Li, M
Published in: ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
January 1, 2022

Nowadays, as more and more systems achieve good performance in traditional voice conversion (VC) tasks, people's attention gradually turns to VC tasks under extreme conditions. In this paper, we propose a novel method for zero-shot voice conversion. We aim to obtain intermediate representations for speaker-content disentanglement of speech to better remove speaker information and get pure content information. Accordingly, our proposed framework contains a module that removes the speaker information from the acoustic feature of the source speaker. Moreover, speaker information control is added to our system to maintain the voice cloning performance. The proposed system is evaluated by subjective and objective metrics. Results show that our proposed system significantly reduces the trade-off problem in zero-shot voice conversion, while it also manages to have high spoofing power to the speaker verification system.

Duke Scholars

Published In

ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings

DOI

ISSN

1520-6149

Publication Date

January 1, 2022

Volume

2022-May

Start / End Page

6567 / 6571
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Zhang, H., Cai, Z., Qin, X., & Li, M. (2022). SIG-VC: A SPEAKER INFORMATION GUIDED ZERO-SHOT VOICE CONVERSION SYSTEM FOR BOTH HUMAN BEINGS AND MACHINES. In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings (Vol. 2022-May, pp. 6567–6571). https://doi.org/10.1109/ICASSP43922.2022.9746048
Zhang, H., Z. Cai, X. Qin, and M. Li. “SIG-VC: A SPEAKER INFORMATION GUIDED ZERO-SHOT VOICE CONVERSION SYSTEM FOR BOTH HUMAN BEINGS AND MACHINES.” In ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2022-May:6567–71, 2022. https://doi.org/10.1109/ICASSP43922.2022.9746048.
Zhang H, Cai Z, Qin X, Li M. SIG-VC: A SPEAKER INFORMATION GUIDED ZERO-SHOT VOICE CONVERSION SYSTEM FOR BOTH HUMAN BEINGS AND MACHINES. In: ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. 2022. p. 6567–71.
Zhang, H., et al. “SIG-VC: A SPEAKER INFORMATION GUIDED ZERO-SHOT VOICE CONVERSION SYSTEM FOR BOTH HUMAN BEINGS AND MACHINES.” ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, vol. 2022-May, 2022, pp. 6567–71. Scopus, doi:10.1109/ICASSP43922.2022.9746048.
Zhang H, Cai Z, Qin X, Li M. SIG-VC: A SPEAKER INFORMATION GUIDED ZERO-SHOT VOICE CONVERSION SYSTEM FOR BOTH HUMAN BEINGS AND MACHINES. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. 2022. p. 6567–6571.

Published In

ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings

DOI

ISSN

1520-6149

Publication Date

January 1, 2022

Volume

2022-May

Start / End Page

6567 / 6571