Theme
Speech and Audio
In a sentence: Speech and audio processing is the intersection of deep learning's "sequence modeling + signal representation" — converting continuous sound waves into feature sequences, then using RNNs and Sequence Modeling or Transformer Architecture to learn "listening" (recognition) and "speaking" (synthesis) — it is one of the earliest areas of deep learning to see large-scale commercial deployment, and is now an important wing of modern multimodal large models.
1. Audio Representations
Audio in a computer is waveform: a one-dimensional array of 16,000~48,000 samples per second. Feeding waveforms directly to neural networks is feasible, but long and redundant — so the industry convention is to transform into the time-frequency domain:
- Short-Time Fourier Transform (STFT): Cut waveforms into frames (~25ms) and perform Fourier transforms, yielding "time × frequency" spectrograms;
- Mel-spectrogram: Compress the frequency axis according to human auditory perception (Mel scale), aligning more closely with hearing — the standard input for TTS and many speech models;
- MFCC: Apply discrete cosine transform to Mel-spectrograms and extract low-dimensional coefficients — the classic ASR feature, now largely superseded by deep features;
- Features vs. end-to-end: Modern end-to-end systems often learn directly from waveforms or raw Mel-spectrograms, letting networks learn features themselves (contrast with feature engineering philosophy in Data and Data Engineering).
2. ASR Evolution: From CTC to Whisper
The core challenge of speech recognition (Automatic Speech Recognition, ASR) is alignment: speech frames (hundreds per second) and text tokens (a few per second) have mismatched lengths and unlabeled alignments. Three generations of solutions:
- HMM-GMM era (pre-deep learning): Hand-crafted acoustic models + pronunciation dictionaries + language models, heavily engineering-intensive;
- CTC (2006, Graves; deepified 2014+): Allows the network to output "repeated + blank" alignment paths, aggregating all possible alignments via a dynamic programming loss — no frame-level annotations needed, the first springboard for deep learning ASR;
- Seq2Seq attention / Transformer: Encode speech into feature sequences, decode text using attention (Listen, Attend and Spell, 2015; thereafter Transformers took over completely).
Whisper (2022, OpenAI) is a paradigm shift for ASR: trained on 680,000 hours of multilingual weakly supervised data (web-scraped, without manual fine-grained annotation) using large-scale Transformers, it achieves recognition, translation, and timestamp alignment across 99 languages — "out of the box" coverage far exceeds all prior systems. It validates that "data scale > architectural elegance" also holds in speech (consistent with the pretraining logic of Large Language Models (LLM)).
3. TTS: From Concatenative Synthesis to Neural
Evolution of Text-to-Speech (TTS):
| System | Year | Type | Characteristics |
|---|---|---|---|
| WaveNet | 2016 | Autoregressive waveform generation (dilated causal conv) | First to surpass concatenative/parametric TTS quality, but per-sample generation was extremely slow |
| Tacotron / Tacotron 2 | 2017/2018 | Seq2Seq text → Mel-spectrogram | End-to-end text-to-speech, producing waveforms via vocoder |
| FastSpeech / parallel TTS | 2019 | Non-autoregressive (length regularization + parallel) | Abandoning per-frame autoregression, 100× speedup |
| VITS | 2021 | End-to-end + VAE + adversarial training | Single model text → waveform, high quality, low latency, the open-source community's mainstay |
| Diffusion TTS (Diff-TTS/Grad-TTS) | 2021 | Diffusion models generate Mel-spectrograms | Quality and naturalness breakthroughs; see the "Diffusion Models and Generative AI" article |
The deeper difficulties of TTS: acoustic non-uniqueness (the same sentence has countless natural pronunciations), prosody and emotion, and long-text stability (repetition, skipped words). In the era of large models, TTS is trending toward "audio language models" (see below).
4. Speech Enhancement and Separation
- Speech enhancement: Recover clean speech from noisy/reverberant audio, essentially regressing to clean targets (often ideal masks/spectra). Core pre-processing module for conference systems, hearing aids, and far-field speech;
- Speech separation: Separate each speaker's voice from a recording with multiple overlapping voices ("the cocktail party problem"). Methods evolved from spectral masks to Permutation Invariant Training (PIT) to recent large-scale separation models (SepFormer, SCNet);
- Commonality: Both are sequence-to-sequence regressions of "noisy input → clean output," with attention/Transformers now the default backbone.
5. Large Audio Models
After 2022, speech and audio accelerated into the "large model" narrative:
- AudioLM (2023, Google): Pretrains audio tokens similarly to BERT/GPT (sound discretized into codec tokens), enabling text-free timbre continuation and cross-lingual transfer;
- Text-to-audio: AudioLDM, Stable Audio, SongGen (2025, ByteDance), etc. use diffusion models/language models for music and sound effect generation (the diffusion path for music generation is covered in Diffusion Models and Generative AI);
- Speech large models: GPT-4o's "voice conversation," various streaming ASR/TTS large models compress "hear-think-speak" into a single model — fully isomorphic to the Multimodal Models "unified modality" path;
- Speech-to-speech: Direct voice-to-voice conversion (preserving emotion and timbre), bypassing text as an intermediate representation.
6. Streaming and Latency Trade-offs
The engineering lifeblood of speech systems is latency:
- Streaming: Produce results while listening; both ASR and TTS must "depend only on already-arrived input," limiting attention's "look-ahead" capability — streaming systems often have to use unidirectional/causal structures (RNNs still have a role here; see the "RNNs and Sequence Modeling" article);
- Offline: Process the full sentence, using bidirectional attention for higher quality (Whisper falls in this category);
- Trade-off strategies: Chunked attention, "fast-first, slow-second" two-pass recognition, and inference acceleration (KV cache, quantization; see the "Large Language Models (LLM)" article) are all engineering countermeasures;
- Beyond WER (word error rate) / MOS (naturalness subjective score), Real-Time Factor (RTF, the time needed to process 1 second of audio) is the key metric for service design.
Engineering Notes
Speech system evaluation is stricter than it seems: the same speaker's fast/slow pace, accent, and noise environment can make model performance vary dramatically. Matching training/test set distributions takes priority over chasing WER numbers; test set design methodology is covered in Deep Learning Evaluation and Experimentation.
7. Trade-offs
- Streaming vs quality: Real-time experiences (meetings, customer service) require streaming; if waiting is fine (subtitles, batch transcription), use offline for better quality;
- Autoregressive vs parallel: Autoregressive TTS is natural but slow; non-autoregressive/diffusion is faster but requires additional mechanisms for stability;
- Specialized vs general-purpose: Large models like Whisper generalize strongly but are heavy and high-latency; lightweight specialized models (on-device ASR) are fast but narrow-scoped — see MLOps and Model Deployment for on-device deployment considerations;
- Features vs end-to-end: Hand-crafted features like MFCCs are cheap and integrate with traditional signal processing; Mel-spectrograms/raw waveforms have higher end-to-end ceilings but demand more data and compute.
Further Reading
- RNNs and Sequence Modeling — Recurrent structures in streaming speech
- Transformer Architecture — The backbone of Whisper/speech large models
- Multimodal Models — How speech enters unified multimodal systems
- Diffusion Models and Generative AI — Diffusion TTS and audio generation
- Large Language Models (LLM) — The paradigm behind speech large models
- Datasets and Tool Archives — LibriSpeech and other speech datasets and tools
References
- Graves et al. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks (ICML 2006)
- Chan et al. Listen, Attend and Spell (ICASSP 2016)
- Radford et al. Robust Speech Recognition via Large-Scale Weak Supervision (Whisper) (ICML 2023)
- van den Oord et al. WaveNet: A Generative Model for Raw Audio (2016)
- Wang et al. Tacotron: Towards End-to-End Speech Synthesis (Interspeech 2017)
- Kim et al. Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech (VITS) (ICML 2021)
- Borsos et al. AudioLM: a Language Modeling Approach to Audio Generation (IEEE/ACM TASLP 2023)
- Panayotov et al. LibriSpeech: an ASR corpus based on public domain audio books (ICASSP 2015)