WavTokenizer: An Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Abstract
Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1) extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2) improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The code is available at https://github.com/jishengpeng/WavTokenizer.
Cite
Text
Ji et al. "WavTokenizer: An Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling." International Conference on Learning Representations, 2025.Markdown
[Ji et al. "WavTokenizer: An Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling." International Conference on Learning Representations, 2025.](https://mlanthology.org/iclr/2025/ji2025iclr-wavtokenizer/)BibTeX
@inproceedings{ji2025iclr-wavtokenizer,
title = {{WavTokenizer: An Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling}},
author = {Ji, Shengpeng and Jiang, Ziyue and Wang, Wen and Chen, Yifu and Fang, Minghui and Zuo, Jialong and Yang, Qian and Cheng, Xize and Wang, Zehan and Li, Ruiqi and Zhang, Ziang and Yang, Xiaoda and Huang, Rongjie and Jiang, Yidi and Chen, Qian and Zheng, Siqi and Zhao, Zhou},
booktitle = {International Conference on Learning Representations},
year = {2025},
url = {https://mlanthology.org/iclr/2025/ji2025iclr-wavtokenizer/}
}