Skip to content

Speech and Audio

Quick overview Speech and audio are among the earliest and most deeply applied areas of deep learning. This article covers representations from waveforms/MFCCs/Mel-spectrograms, traces ASR evolution from CTC to Seq2Seq to Whisper's large-scale weak supervision, TTS from Tacotron/WaveNet to VITS and diffusion TTS, and introduces speech enhancement, separation, large audio models, and streaming latency engineering trade-offs.

Speech and Audio ​

In a sentence: Speech and audio processing is the intersection of deep learning's "sequence modeling + signal representation" — converting continuous sound waves into feature sequences, then using RNNs and Sequence Modeling or Transformer Architecture to learn "listening" (recognition) and "speaking" (synthesis) — it is one of the earliest areas of deep learning to see large-scale commercial deployment, and is now an important wing of modern multimodal large models.

1. Audio Representations ​

Audio in a computer is waveform: a one-dimensional array of 16,000~48,000 samples per second. Feeding waveforms directly to neural networks is feasible, but long and redundant — so the industry convention is to transform into the time-frequency domain:

  • Short-Time Fourier Transform (STFT): Cut waveforms into frames (~25ms) and perform Fourier transforms, yielding "time × frequency" spectrograms;
  • Mel-spectrogram: Compress the frequency axis according to human auditory perception (Mel scale), aligning more closely with hearing — the standard input for TTS and many speech models;
  • MFCC: Apply discrete cosine transform to Mel-spectrograms and extract low-dimensional coefficients — the classic ASR feature, now largely superseded by deep features;
  • Features vs. end-to-end: Modern end-to-end systems often learn directly from waveforms or raw Mel-spectrograms, letting networks learn features themselves (contrast with feature engineering philosophy in Data and Data Engineering).

2. ASR Evolution: From CTC to Whisper ​

The core challenge of speech recognition (Automatic Speech Recognition, ASR) is alignment: speech frames (hundreds per second) and text tokens (a few per second) have mismatched lengths and unlabeled alignments. Three generations of solutions:

  1. HMM-GMM era (pre-deep learning): Hand-crafted acoustic models + pronunciation dictionaries + language models, heavily engineering-intensive;
  2. CTC (2006, Graves; deepified 2014+): Allows the network to output "repeated + blank" alignment paths, aggregating all possible alignments via a dynamic programming loss — no frame-level annotations needed, the first springboard for deep learning ASR;
  3. Seq2Seq attention / Transformer: Encode speech into feature sequences, decode text using attention (Listen, Attend and Spell, 2015; thereafter Transformers took over completely).

Whisper (2022, OpenAI) is a paradigm shift for ASR: trained on 680,000 hours of multilingual weakly supervised data (web-scraped, without manual fine-grained annotation) using large-scale Transformers, it achieves recognition, translation, and timestamp alignment across 99 languages — "out of the box" coverage far exceeds all prior systems. It validates that "data scale > architectural elegance" also holds in speech (consistent with the pretraining logic of Large Language Models (LLM)).

3. TTS: From Concatenative Synthesis to Neural ​

Evolution of Text-to-Speech (TTS):

SystemYearTypeCharacteristics
WaveNet2016Autoregressive waveform generation (dilated causal conv)First to surpass concatenative/parametric TTS quality, but per-sample generation was extremely slow
Tacotron / Tacotron 22017/2018Seq2Seq text → Mel-spectrogramEnd-to-end text-to-speech, producing waveforms via vocoder
FastSpeech / parallel TTS2019Non-autoregressive (length regularization + parallel)Abandoning per-frame autoregression, 100× speedup
VITS2021End-to-end + VAE + adversarial trainingSingle model text → waveform, high quality, low latency, the open-source community's mainstay
Diffusion TTS (Diff-TTS/Grad-TTS)2021Diffusion models generate Mel-spectrogramsQuality and naturalness breakthroughs; see the "Diffusion Models and Generative AI" article

The deeper difficulties of TTS: acoustic non-uniqueness (the same sentence has countless natural pronunciations), prosody and emotion, and long-text stability (repetition, skipped words). In the era of large models, TTS is trending toward "audio language models" (see below).

4. Speech Enhancement and Separation ​

  • Speech enhancement: Recover clean speech from noisy/reverberant audio, essentially regressing to clean targets (often ideal masks/spectra). Core pre-processing module for conference systems, hearing aids, and far-field speech;
  • Speech separation: Separate each speaker's voice from a recording with multiple overlapping voices ("the cocktail party problem"). Methods evolved from spectral masks to Permutation Invariant Training (PIT) to recent large-scale separation models (SepFormer, SCNet);
  • Commonality: Both are sequence-to-sequence regressions of "noisy input → clean output," with attention/Transformers now the default backbone.

5. Large Audio Models ​

After 2022, speech and audio accelerated into the "large model" narrative:

  • AudioLM (2023, Google): Pretrains audio tokens similarly to BERT/GPT (sound discretized into codec tokens), enabling text-free timbre continuation and cross-lingual transfer;
  • Text-to-audio: AudioLDM, Stable Audio, SongGen (2025, ByteDance), etc. use diffusion models/language models for music and sound effect generation (the diffusion path for music generation is covered in Diffusion Models and Generative AI);
  • Speech large models: GPT-4o's "voice conversation," various streaming ASR/TTS large models compress "hear-think-speak" into a single model — fully isomorphic to the Multimodal Models "unified modality" path;
  • Speech-to-speech: Direct voice-to-voice conversion (preserving emotion and timbre), bypassing text as an intermediate representation.

6. Streaming and Latency Trade-offs ​

The engineering lifeblood of speech systems is latency:

  • Streaming: Produce results while listening; both ASR and TTS must "depend only on already-arrived input," limiting attention's "look-ahead" capability — streaming systems often have to use unidirectional/causal structures (RNNs still have a role here; see the "RNNs and Sequence Modeling" article);
  • Offline: Process the full sentence, using bidirectional attention for higher quality (Whisper falls in this category);
  • Trade-off strategies: Chunked attention, "fast-first, slow-second" two-pass recognition, and inference acceleration (KV cache, quantization; see the "Large Language Models (LLM)" article) are all engineering countermeasures;
  • Beyond WER (word error rate) / MOS (naturalness subjective score), Real-Time Factor (RTF, the time needed to process 1 second of audio) is the key metric for service design.

Engineering Notes

Speech system evaluation is stricter than it seems: the same speaker's fast/slow pace, accent, and noise environment can make model performance vary dramatically. Matching training/test set distributions takes priority over chasing WER numbers; test set design methodology is covered in Deep Learning Evaluation and Experimentation.

7. Trade-offs ​

  • Streaming vs quality: Real-time experiences (meetings, customer service) require streaming; if waiting is fine (subtitles, batch transcription), use offline for better quality;
  • Autoregressive vs parallel: Autoregressive TTS is natural but slow; non-autoregressive/diffusion is faster but requires additional mechanisms for stability;
  • Specialized vs general-purpose: Large models like Whisper generalize strongly but are heavy and high-latency; lightweight specialized models (on-device ASR) are fast but narrow-scoped — see MLOps and Model Deployment for on-device deployment considerations;
  • Features vs end-to-end: Hand-crafted features like MFCCs are cheap and integrate with traditional signal processing; Mel-spectrograms/raw waveforms have higher end-to-end ceilings but demand more data and compute.

Further Reading ​

References ​