Skip to main content

Xiaoxiao Miao

Assistant Professor of Computer Science at Duke Kunshan University
DKU Faculty

Scholarly Works - Journal articles


Privacy attacks on voice anonymization systems: Overview and key findings from the First VoicePrivacy Attacker Challenge

Journal article Computer Speech and Language · February 1, 2027 The First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge that focuses on evaluating attacker systems against a set of voice anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation data ... Full text Cite

The third VoicePrivacy challenge: Preserving emotional expressiveness and linguistic content in voice anonymization

Journal article Computer Speech and Language · October 1, 2026 We present results and analyses from the third VoicePrivacy Challenge held in 2024, which focuses on advancing voice anonymization technologies. The task was to develop a voice anonymization system for speech data that conceals a speaker’s voice identity w ... Full text Cite

Spectrogram Features for Audio and Speech Analysis

Journal article Applied Sciences Switzerland · January 1, 2026 Featured Application: Spectrogram-based input features have become the most popular choice for deep learning models that classify audio and speech, yet there are many settings related to resolution and representation type. This article surveys those choice ... Full text Cite

Adapting general disentanglement-based speaker anonymization for enhanced emotion preservation

Journal article Computer Speech and Language · November 1, 2025 A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, ... Full text Cite

A Benchmark for Multi-Speaker Anonymization

Journal article IEEE Transactions on Information Forensics and Security · January 1, 2025 Privacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus particularly on single-speaker scenarios. However, they l ... Full text Cite

Joint speaker encoder and neural back-end model for fully end-to-end automatic speaker verification with multiple enrollment utterances

Journal article Computer Speech and Language · June 1, 2024 Conventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and a back-end model such as statistics-based probabilistic linear discriminant ... Full text Cite

VoicePAT: An Efficient Open-Source Evaluation Toolkit for Voice Privacy Research

Journal article IEEE Open Journal of Signal Processing · January 1, 2024 Speaker anonymization is the task of modifying a speech recording such that the original speaker cannot be identified anymore. Since the first Voice Privacy Challenge in 2020, along with the release of a framework, the popularity of this research topic is ... Full text Cite

The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation

Journal article IEEE ACM Transactions on Audio Speech and Language Processing · January 1, 2024 —The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datase ... Full text Cite

Speaker Anonymization Using Orthogonal Householder Neural Network

Journal article IEEE ACM Transactions on Audio Speech and Language Processing · January 1, 2023 Speaker anonymization aims to conceal a speaker's identity while preserving content information in speech. Current mainstream neural-network speaker anonymization systems disentangle speech into prosody-related, content, and speaker representations. The sp ... Full text Cite

GuidedMix: An on-the-fly data augmentation approach for robust speaker recognition system

Journal article Electronics Letters · January 1, 2022 Data augmentation is an essential technique for building a high-robustness speaker recognition system. this letter proposes a novel on-the-fly data augmentation strategy called GuidedMix. It significantly increases augmented data fidelity by mixing the spe ... Full text Cite

Variance Normalised Features for Language and Dialect Discrimination

Journal article Circuits Systems and Signal Processing · July 1, 2021 This paper proposes novel features for automated language and dialect identification that aim to improve discriminative power by ensuring that each element of the feature vector has a normalised contribution to inter-class variance. The method firstly comp ... Full text Cite

D-MONA: A dilated mixed-order non-local attention network for speaker and language recognition.

Journal article Neural networks : the official journal of the International Neural Network Society · July 2021 Attention-based convolutional neural network (CNN) models are increasingly being adopted for speaker and language recognition (SR/LR) tasks. These include time, frequency, spatial and channel attention, which can focus on useful time frames, frequency band ... Full text Cite

Cross-domain speaker recognition using domain adversarial siamese network with a domain discriminator

Journal article Electronics Letters · July 9, 2020 With the widespread use of automatic speaker recognition in realistic world, it suffers a lot when there is a domain mismatch, including channel, language, distance etc. Recent research studies have introduced the adversarial-learning mechanism into deep n ... Full text Cite

A New Time–Frequency Attention Tensor Network for Language Identification

Journal article Circuits Systems and Signal Processing · May 1, 2020 In this paper, we aim to improve traditional DNN x-vector language identification performance by employing wide residual networks (WRN) as a powerful feature extractor which we combine with a novel frequency attention network. Compared with conventional ti ... Full text Cite

Denoising Autoencoder-Based Language Feature Compensation

Journal article Jisuanji Yanjiu Yu Fazhan Computer Research and Development · May 1, 2019 Language identification (LID) accuracy is often significantly reduced when the duration of the test data and the training data are mismatched. This paper proposes a method to compensate language features using a denoising autoencoder (DAE). Use of denoisin ... Full text Cite

Expanding the length of short utterances for short-duration language recognition

Journal article Qinghua Daxue Xuebao Journal of Tsinghua University · March 1, 2018 The language recognition (LR) accuracy is often significantly reduced when the test utterance duration is as short as 10 s or less. This paper describes a method to extend the utterance length using time-scale modification (TSM) which changes the speech ra ... Full text Cite