Skip to main content

Xiaoxiao Miao

Assistant Professor of Computer Science at Duke Kunshan University
DKU Faculty

Scholarly Works - Conferences


CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation

Conference 2026 IEEE International Conference on Human Machine Systems Ichms 2026 · January 1, 2026 Instruction-guided text-to-speech (TTS) research has reached a maturity level where excellent speech generation quality is possible on demand, yet two coupled biases persist in reducing perceived quality: accent bias, where models default towards dominant ... Full text Cite

The First VoicePrivacy Attacker Challenge

Conference ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · January 1, 2025 The First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge which focuses on evaluating attacker systems against a set of voice anonymization systems submitted to the VoicePrivacy 2024 Challenge. Training, development, and evaluation dat ... Full text Cite

Automated evaluation of children's speech fluency for low-resource languages

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2025 Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a system to automatically assess fluency by combining a fine-tuned multilingual ... Full text Cite

LSPnet: an ultra-low bitrate hybrid neural codec

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2025 This paper presents an ultra-low bitrate speech codec that achieves high-fidelity speech coding at 1.2kbps while maintaining low computational complexity. Building upon the LPCNet framework, combined with a parametric encoder, we introduce several key impr ... Full text Cite

Speech Emotion Recognition Via Entropy-Aware Score Selection

Conference 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference Apsipa ASC 2025 · January 1, 2025 In this paper, we propose multimodal framework for speech emotion recognition that leverages entropy-aware score selection to combine speech and textual predictions. The proposed method integrates a primary pipeline that consists of an acoustic model based ... Full text Cite

SegReConcat: A Data Augmentation Method for Voice Anonymization Attack

Conference 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference Apsipa ASC 2025 · January 1, 2025 Anonymization of voice seeks to conceal the identity of the speaker while maintaining the utility of speech data. However, residual speaker cues often persist, which pose privacy risks. We propose SegReConcat, a data augmentation method for attacker-side e ... Full text Cite

Exploring Machine Learning and Language Models for Multimodal Depression Detection

Conference 2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference Apsipa ASC 2025 · January 1, 2025 This paper presents our approach to the first Multimodal Personality-Aware Depression Detection Challenge, focusing on multimodal depression detection using machine learning and deep learning models. We explore and compare the performance of XGBoost, trans ... Full text Cite

SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization

Conference Asru 2025 2025 IEEE Automatic Speech Recognition and Understanding Workshop · January 1, 2025 Voice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-em ... Full text Cite

SecureSpeech: Prompt-based Speaker and Content Protection

Conference 2025 IEEE International Joint Conference on Biometrics Ijcb 2025 · January 1, 2025 Given the increasing privacy concerns from identity theft and the re-identification of speakers through content in the speech field, this paper proposes a prompt-based speech generation pipeline that ensures dual anonymization of both speaker identity and ... Full text Cite

Target Speaker Extraction with Curriculum Learning

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2024 This paper presents a novel approach to target speaker extraction (TSE) using Curriculum Learning (CL) techniques, addressing the challenge of distinguishing a target speaker's voice from a mixture containing interfering speakers. For efficient training, w ... Full text Cite

Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches

Conference Proceedings of 2024 IEEE Spoken Language Technology Workshop Slt 2024 · January 1, 2024 In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel mismatch, and domain mismatch. Traditional automatic speaker verification (ASV) s ... Full text Cite

SYNVOX2: TOWARDS A PRIVACY-FRIENDLY VOXCELEB2 DATASET

Conference ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · January 1, 2024 The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using ... Full text Cite

Instructsing: High-Fidelity Singing Voice Generation Via Instructing Yourself

Conference Proceedings of 2024 IEEE Spoken Language Technology Workshop Slt 2024 · January 1, 2024 It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster compared with other ... Full text Cite

Improving Generalization Ability of Countermeasures for New Mismatch Scenario by Combining Multiple Advanced Regularization Terms

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2023 The ability of countermeasure models to generalize from seen speech synthesis methods to unseen ones has been investigated in the ASVspoof challenge. However, a new mismatch scenario in which fake audio may be generated from real audio with unseen genres h ... Full text Cite

Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline

Conference ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · January 1, 2023 The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the spe ... Full text Cite

Analyzing Language-Independent Speaker Anonymization Framework under Unseen Conditions

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2022 In our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models. Although the system can anonymize speech data of any language, the anonymization was imperfect, and the speech content of the an ... Full text Cite

ATTENTION BACK-END FOR AUTOMATIC SPEAKER VERIFICATION WITH MULTIPLE ENROLLMENT UTTERANCES

Conference ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · January 1, 2022 Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propo ... Full text Cite

Adaptive margin circle loss for speaker verification

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2021 Deep-Neural-Network (DNN) based speaker verification systems use the angular softmax loss with margin penalties to enhance the intra-class compactness of speaker embeddings, which achieved remarkable performance. In this paper, we propose a novel angular l ... Full text Cite

A new time-frequency attention mechanism for TDNN and CNN-LSTM-TDNN, with application to language identification

Conference Proceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2019 In this paper, we aim to improve traditional DNN x-vector language identification (LID) performance by employing Convolutional and Long Short Term Memory-Recurrent (CLSTM) Neural Networks, as they can strengthen feature extraction and capture longer tempor ... Full text Cite

Improved Conditional Generative Adversarial Net Classification for Spoken Language Recognition

Conference 2018 IEEE Spoken Language Technology Workshop Slt 2018 Proceedings · July 2, 2018 Recent research on generative adversarial nets (GAN) for language identification (LID) has shown promising results. In this paper, we further exploit the latent abilities of GAN networks to firstly combine them with deep neural network (DNN)-based i-vector ... Full text Cite