Skip to main content

Spectrogram Features for Audio and Speech Analysis

Journal articles  - Review
McLoughlin, I; Pham, L; Song, Y; Miao, X; Phan, H; Cai, P; Gu, Q; Nan, J; Song, H; Soh, D
Published in: Applied Sciences Switzerland
January 1, 2026

Featured Application: Spectrogram-based input features have become the most popular choice for deep learning models that classify audio and speech, yet there are many settings related to resolution and representation type. This article surveys those choices and discusses their suitability for different application areas. Spectrogram-based representations have grown to dominate the feature space for deep learning audio analysis systems, and are often adopted for speech analysis also. Initially, the primary motivation behind spectrogram-based representations was their ability to present sound as a two-dimensional signal in the time–frequency plane, which not only provides an interpretable physical basis for analysing sound, but also unlocks the use of a range of machine learning techniques such as convolutional neural networks, which had been developed for image processing. A spectrogram is a matrix characterised by the resolution and span of its dimensions, as well as by the representation and scaling of each element. Many possibilities for these three characteristics have been explored by researchers across numerous application areas, with different settings showing affinity for various tasks. This paper reviews the use of spectrogram-based representations and surveys the state-of-the-art to question how front-end feature representation choice allies with back-end classifier architecture for different tasks.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Applied Sciences Switzerland

DOI

EISSN

2076-3417

Publication Date

January 1, 2026

Volume

16

Issue

2
 

Citation

APA
Chicago
ICMJE
MLA
NLM
McLoughlin, I., Pham, L., Song, Y., Miao, X., Phan, H., Cai, P., … Soh, D. (2026). Spectrogram Features for Audio and Speech Analysis. Applied Sciences Switzerland, 16(2). https://doi.org/10.3390/app16020572
McLoughlin, I., L. Pham, Y. Song, X. Miao, H. Phan, P. Cai, Q. Gu, J. Nan, H. Song, and D. Soh. “Spectrogram Features for Audio and Speech Analysis.” Applied Sciences Switzerland 16, no. 2 (January 1, 2026). https://doi.org/10.3390/app16020572.
McLoughlin I, Pham L, Song Y, Miao X, Phan H, Cai P, et al. Spectrogram Features for Audio and Speech Analysis. Applied Sciences Switzerland. 2026 Jan 1;16(2).
McLoughlin, I., et al. “Spectrogram Features for Audio and Speech Analysis.” Applied Sciences Switzerland, vol. 16, no. 2, Jan. 2026. Scopus, doi:10.3390/app16020572.
McLoughlin I, Pham L, Song Y, Miao X, Phan H, Cai P, Gu Q, Nan J, Song H, Soh D. Spectrogram Features for Audio and Speech Analysis. Applied Sciences Switzerland. 2026 Jan 1;16(2).

Published In

Applied Sciences Switzerland

DOI

EISSN

2076-3417

Publication Date

January 1, 2026

Volume

16

Issue

2