01/10/2026
01/10/2026
01/10/2026
NeuCodec in Transformers 5.17
NeuCodec in Transformers 5.17
NeuCodec in Transformers 5.17


Simplified 3D illustration of FSQ. NeuCodec uses eight dimensions.
Simplified 3D illustration of FSQ. NeuCodec uses eight dimensions.
NeuCodec in Transformers 5.17
NeuTTS-2E: on-device emotional TTS,
built for the sentences that don't help it
NeuCodec in Transformers 5.17
NeuCodec is now integrated into Hugging Face Transformers 5.17.
Neucodec, our neural audio codec, provides the audio representation for our on-device speech models. Having been widely adopted by the community, Neucodec has amassed over 2.3 million downloads in the past year alone.
NeuCodec uses finite scalar quantization (FSQ) to turn the encoder’s continuous outputs into discrete audio codes. FSQ bounds and rounds each value independently to a fixed set of levels. Together, these rounded values identify a point on a grid, as illustrated in the above graphic.
What makes NeuCodec unique is its use of a single codebook alongside a low token rate: 1 second of audio is translated into 50 tokens. As an example, “Hello, how are you” can become something akin to [12, 35, 14, …, 62300, 38457].
To continue to drive adoption and make NeuCodec accessible through the tools developers already use, we collaborated with Eric and the team at Hugging Face to integrate it directly into their transformers library. This has a number of advantages:
For developers, it means not having to manage multiple packages with conflicting dependencies in their environment, as well as having the option to use it straight out of the box with the common
AutoModelAPI. Given coding agents are so widely used now, it could also reduce any friction from out-of-date training data by letting them work with the familiar Transformers interface, rather than requiring knowledge of a separate package’s API.For us over at Neuphonic, we know that a version of our model is now integrated and maintained in one of the largest packages in the AI ecosystem, which takes advantage of the robust CI/CD and skilled maintainers of the project. We’re fortunate to have the support of Hugging Face’s experienced maintainers and testing infrastructure as we continue to maintain the integration and address package updates or dependency changes.
We worked closely with Hugging Face’s audio team in planning the pull request. We began with a common base and built our implementation on top. We used Transformers’ modular API to express those differences, inheriting components from an existing model and overriding only the parts that needed to change. The library’s tooling then generates standalone implementation files from that modular source, allowing us to build on a base model while keeping NeuCodec’s changes explicit in the source we maintain.
There were a couple of additional technical details to consider in our implementation.
First, batch preprocessing for the semantic encoder’s mel-spectrograms wasn’t particularly straightforward because of torchaudio.compliance.kaldi.fbank, which does not support batched inputs, and the need to preserve consistency with the original model’s outputs. The maintainers agreed to leave this as a possible improvement for NeuCodec, rather than a requirement for the integration.
Second, NeuCodec’s sample rates also raised a related interface question. It accepts audio at 16 kHz and reconstructs upsampled audio at 24 kHz, so a single sampling_rate field could describe either the encoder’s input or the decoder’s output. We made both explicit in the model configuration as input_sampling_rate and output_sampling_rate. The feature extractor’s sampling_rate continues to refer to the input. This gives each setting a defined role while accommodating a model whose input and output rates differ.
To get started with the integration, install transformers 5.17 with pip:
NeuCodec is now integrated into Hugging Face Transformers 5.17.
Neucodec, our neural audio codec, provides the audio representation for our on-device speech models. Having been widely adopted by the community, Neucodec has amassed over 2.3 million downloads in the past year alone.
NeuCodec uses finite scalar quantization (FSQ) to turn the encoder’s continuous outputs into discrete audio codes. FSQ bounds and rounds each value independently to a fixed set of levels. Together, these rounded values identify a point on a grid, as illustrated in the above graphic.
What makes NeuCodec unique is its use of a single codebook alongside a low token rate: 1 second of audio is translated into 50 tokens. As an example, “Hello, how are you” can become something akin to [12, 35, 14, …, 62300, 38457].
To continue to drive adoption and make NeuCodec accessible through the tools developers already use, we collaborated with Eric and the team at Hugging Face to integrate it directly into their transformers library. This has a number of advantages:
For developers, it means not having to manage multiple packages with conflicting dependencies in their environment, as well as having the option to use it straight out of the box with the common
AutoModelAPI. Given coding agents are so widely used now, it could also reduce any friction from out-of-date training data by letting them work with the familiar Transformers interface, rather than requiring knowledge of a separate package’s API.For us over at Neuphonic, we know that a version of our model is now integrated and maintained in one of the largest packages in the AI ecosystem, which takes advantage of the robust CI/CD and skilled maintainers of the project. We’re fortunate to have the support of Hugging Face’s experienced maintainers and testing infrastructure as we continue to maintain the integration and address package updates or dependency changes.
We worked closely with Hugging Face’s audio team in planning the pull request. We began with a common base and built our implementation on top. We used Transformers’ modular API to express those differences, inheriting components from an existing model and overriding only the parts that needed to change. The library’s tooling then generates standalone implementation files from that modular source, allowing us to build on a base model while keeping NeuCodec’s changes explicit in the source we maintain.
There were a couple of additional technical details to consider in our implementation.
First, batch preprocessing for the semantic encoder’s mel-spectrograms wasn’t particularly straightforward because of torchaudio.compliance.kaldi.fbank, which does not support batched inputs, and the need to preserve consistency with the original model’s outputs. The maintainers agreed to leave this as a possible improvement for NeuCodec, rather than a requirement for the integration.
Second, NeuCodec’s sample rates also raised a related interface question. It accepts audio at 16 kHz and reconstructs upsampled audio at 24 kHz, so a single sampling_rate field could describe either the encoder’s input or the decoder’s output. We made both explicit in the model configuration as input_sampling_rate and output_sampling_rate. The feature extractor’s sampling_rate continues to refer to the input. This gives each setting a defined role while accommodating a model whose input and output rates differ.
To get started with the integration, install transformers 5.17 with pip:
NeuCodec is now integrated into Hugging Face Transformers 5.17.
Neucodec, our neural audio codec, provides the audio representation for our on-device speech models. Having been widely adopted by the community, Neucodec has amassed over 2.3 million downloads in the past year alone.
NeuCodec uses finite scalar quantization (FSQ) to turn the encoder’s continuous outputs into discrete audio codes. FSQ bounds and rounds each value independently to a fixed set of levels. Together, these rounded values identify a point on a grid, as illustrated in the above graphic.
What makes NeuCodec unique is its use of a single codebook alongside a low token rate: 1 second of audio is translated into 50 tokens. As an example, “Hello, how are you” can become something akin to [12, 35, 14, …, 62300, 38457].
To continue to drive adoption and make NeuCodec accessible through the tools developers already use, we collaborated with Eric and the team at Hugging Face to integrate it directly into their transformers library. This has a number of advantages:
For developers, it means not having to manage multiple packages with conflicting dependencies in their environment, as well as having the option to use it straight out of the box with the common
AutoModelAPI. Given coding agents are so widely used now, it could also reduce any friction from out-of-date training data by letting them work with the familiar Transformers interface, rather than requiring knowledge of a separate package’s API.For us over at Neuphonic, we know that a version of our model is now integrated and maintained in one of the largest packages in the AI ecosystem, which takes advantage of the robust CI/CD and skilled maintainers of the project. We’re fortunate to have the support of Hugging Face’s experienced maintainers and testing infrastructure as we continue to maintain the integration and address package updates or dependency changes.
We worked closely with Hugging Face’s audio team in planning the pull request. We began with a common base and built our implementation on top. We used Transformers’ modular API to express those differences, inheriting components from an existing model and overriding only the parts that needed to change. The library’s tooling then generates standalone implementation files from that modular source, allowing us to build on a base model while keeping NeuCodec’s changes explicit in the source we maintain.
There were a couple of additional technical details to consider in our implementation.
First, batch preprocessing for the semantic encoder’s mel-spectrograms wasn’t particularly straightforward because of torchaudio.compliance.kaldi.fbank, which does not support batched inputs, and the need to preserve consistency with the original model’s outputs. The maintainers agreed to leave this as a possible improvement for NeuCodec, rather than a requirement for the integration.
Second, NeuCodec’s sample rates also raised a related interface question. It accepts audio at 16 kHz and reconstructs upsampled audio at 24 kHz, so a single sampling_rate field could describe either the encoder’s input or the decoder’s output. We made both explicit in the model configuration as input_sampling_rate and output_sampling_rate. The feature extractor’s sampling_rate continues to refer to the input. This gives each setting a defined role while accommodating a model whose input and output rates differ.
To get started with the integration, install transformers 5.17 with pip:
pip install transformers==5.17And use the following code snippet to get started with neucodec using the AutoModel API:
And use the following code snippet to get started with neucodec using the AutoModel API:
from datasets import Audio, load_dataset
from transformers import AutoFeatureExtractor, AutoModel
model_id = "neuphonic/neucodec"
model = AutoModel.from_pretrained(model_id, device_map="auto")
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
dataset = load_dataset("hf-internal-testing/librispeech_asr_dummy", "clean", split="validation")
dataset = dataset.cast_column("audio", Audio(sampling_rate=feature_extractor.sampling_rate))
audio = dataset[0]["audio"]["array"]
inputs = feature_extractor(audio=audio, sampling_rate=feature_extractor.sampling_rate, return_tensors="pt").to(
model.device, model.dtype
)
print("Input waveform shape:", inputs["input_values"].shape)
# Input waveform shape: torch.Size([1, 1, 93760])
# encoder and decoder
audio_codes = model.encode(**inputs).audio_codes
print("Audio codes shape:", audio_codes.shape)
# Audio codes shape: torch.Size([1, 1, 293])
audio_values = model.decode(audio_codes).audio_values
print("Audio values shape:", audio_values.shape)
# Audio values shape: torch.Size([1, 1, 93760])
# Equivalently, you can do encoding and decoding in one step
model_output = model(**inputs)
audio_codes = model_output.audio_codes
audio_values = model_output.audio_valuesfrom datasets import Audio, load_dataset
from transformers import AutoFeatureExtractor, AutoModel
model_id = "neuphonic/neucodec"
model = AutoModel.from_pretrained(model_id, device_map="auto")
feature_extractor = AutoFeatureExtractor.from_pretrained(model_id)
dataset = load_dataset("hf-internal-testing/librispeech_asr_dummy", "clean", split="validation")
dataset = dataset.cast_column("audio", Audio(sampling_rate=feature_extractor.sampling_rate))
audio = dataset[0]["audio"]["array"]
inputs = feature_extractor(audio=audio, sampling_rate=feature_extractor.sampling_rate, return_tensors="pt").to(
model.device, model.dtype
)
print("Input waveform shape:", inputs["input_values"].shape)
# Input waveform shape: torch.Size([1, 1, 93760])
# encoder and decoder
audio_codes = model.encode(**inputs).audio_codes
print("Audio codes shape:", audio_codes.shape)
# Audio codes shape: torch.Size([1, 1, 293])
audio_values = model.decode(audio_codes).audio_values
print("Audio values shape:", audio_values.shape)
# Audio values shape: torch.Size([1, 1, 93760])
# Equivalently, you can do encoding and decoding in one step
model_output = model(**inputs)
audio_codes = model_output.audio_codes
audio_values = model_output.audio_valuesWe’re grateful to the Hugging Face team for their guidance throughout the integration. In particular, thank you to Eric Bezzam for his detailed reviews and work on this PR!
We’re grateful to the Hugging Face team for their guidance throughout the integration. In particular, thank you to Eric Bezzam for his detailed reviews and work on this PR!
We’re grateful to the Hugging Face team for their guidance throughout the integration. In particular, thank you to Eric Bezzam for his detailed reviews and work on this PR!
#neucodec
#neucodec
#productlaunch
#huggingface
#huggingface
#huggingface
#audioAI
#audioAI
#audioAI
Bring Neuphonic into your product.
Building something with voice AI? Talk to us about your use case,
deployment requirements, and how Neuphonic can fit into your stack.
Building something with voice AI? Talk to us about your use case,
deployment requirements, and how Neuphonic can fit into your stack.
Neuphonic © 2026

Neuphonic © 2026

Neuphonic © 2026

