Scholars@Duke publication: "hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World

"hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World

Publication , Conference

Wenger, E; Bronckers, M; Cianfarani, C; Cryan, J; Sha, A; Zheng, H; Zhao, BY

Published in: Proceedings of the ACM Conference on Computer and Communications Security

November 13, 2021

Advances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of powerful attacks against both humans and software systems (aka machines). This paper documents efforts and findings from a comprehensive experimental study on the impact of deep-learning based speech synthesis attacks on both human listeners and machines such as speaker recognition and voice-signin systems. We find that both humans and machines can be reliably fooled by synthetic speech, and that existing defenses against synthesized speech fall short. These findings highlight the need to raise awareness and develop new protections against synthetic speech for both humans and machines.

Duke Scholars

Author Emily Wenger Electrical and Computer Engineering

Published In

Proceedings of the ACM Conference on Computer and Communications Security

DOI

10.1145/3460120.3484742

ISSN

1543-7221

Publication Date

November 13, 2021

Start / End Page

235 / 251

Citation

APA

Chicago

ICMJE

MLA

NLM

Wenger, E., Bronckers, M., Cianfarani, C., Cryan, J., Sha, A., Zheng, H., & Zhao, B. Y. (2021). "hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World. In Proceedings of the ACM Conference on Computer and Communications Security (pp. 235–251). https://doi.org/10.1145/3460120.3484742

Wenger, E., M. Bronckers, C. Cianfarani, J. Cryan, A. Sha, H. Zheng, and B. Y. Zhao. “"hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World.” In Proceedings of the ACM Conference on Computer and Communications Security, 235–51, 2021. https://doi.org/10.1145/3460120.3484742.

Wenger E, Bronckers M, Cianfarani C, Cryan J, Sha A, Zheng H, et al. "hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World. In: Proceedings of the ACM Conference on Computer and Communications Security. 2021. p. 235–51.

Wenger, E., et al. “"hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World.” Proceedings of the ACM Conference on Computer and Communications Security, 2021, pp. 235–51. Scopus, doi:10.1145/3460120.3484742.

Wenger E, Bronckers M, Cianfarani C, Cryan J, Sha A, Zheng H, Zhao BY. "hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World. Proceedings of the ACM Conference on Computer and Communications Security. 2021. p. 235–251.

Published In

Proceedings of the ACM Conference on Computer and Communications Security

DOI

10.1145/3460120.3484742

ISSN

1543-7221

Publication Date

November 13, 2021

Start / End Page

235 / 251