Skip to main content
Journal cover image

Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems

Publication ,  Journal Article
Das, A; Soni, P; Huang, MC; Lin, F; Xu, W
Published in: Smart Health
June 1, 2024

Speech recognition using EEG signals captured during covert (imagined) speech has garnered substantial interest in Brain–Computer Interface (BCI) research. While the concept holds promise, current implementations must improve performance compared to established Automatic Speech Recognition (ASR) methods using audio. An area often underestimated in previous studies is the potential of EEG utilization during overt speech. Integrating overt EEG signals with speech data by leveraging advancements in deep learning presents significant potential to enhance the efficacy of these systems. This integration proves particularly advantageous in noisy environments and for individuals with speech impairments—challenges even conventional ASR techniques struggle to address effectively. Our investigation delves into this relationship by introducing a novel multimodal model that merges EEG and speech inputs. Our model achieves a multiclass classification accuracy of 95.39%. When subjected to artificial white noise added to the input audio, our model exhibits a notable level of resilience, surpassing the capabilities of models reliant solely on single EEG or audio modalities. The validation process, leveraging the robust techniques of t-SNE and silhouette coefficient, corroborates and solidifies these advancements.

Duke Scholars

Published In

Smart Health

DOI

EISSN

2352-6483

Publication Date

June 1, 2024

Volume

32

Related Subject Headings

  • 46 Information and computing sciences
  • 42 Health sciences
  • 40 Engineering
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Das, A., Soni, P., Huang, M. C., Lin, F., & Xu, W. (2024). Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems. Smart Health, 32. https://doi.org/10.1016/j.smhl.2024.100477
Das, A., P. Soni, M. C. Huang, F. Lin, and W. Xu. “Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems.” Smart Health 32 (June 1, 2024). https://doi.org/10.1016/j.smhl.2024.100477.
Das A, Soni P, Huang MC, Lin F, Xu W. Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems. Smart Health. 2024 Jun 1;32.
Das, A., et al. “Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems.” Smart Health, vol. 32, June 2024. Scopus, doi:10.1016/j.smhl.2024.100477.
Das A, Soni P, Huang MC, Lin F, Xu W. Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems. Smart Health. 2024 Jun 1;32.
Journal cover image

Published In

Smart Health

DOI

EISSN

2352-6483

Publication Date

June 1, 2024

Volume

32

Related Subject Headings

  • 46 Information and computing sciences
  • 42 Health sciences
  • 40 Engineering