Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
Quick Answer
Hugging Face introduces Nunchaku Lite, enabling 4-bit diffusion inference in Diffusers without custom pipelines, achieving 30% speedup and 50% VRAM reduction.
Quick Take
The NVFP4 checkpoints generate 1024x1024 images in 1.7 seconds on RTX 5090, compared to 24 GB VRAM for BF16 models.
Key Points
- Nunchaku Lite simplifies loading 4-bit models in Diffusers with no local compilation.
- SVDQuant quantization improves speed and reduces memory for diffusion transformers.
- NVFP4 checkpoints require NVIDIA Blackwell GPUs, offering significant performance gains.
- The toolkit allows users to quantize and publish their own models easily.
- Nunchaku Lite achieves around 30% speedup with VRAM reduction compared to standard models.
DeepSignal Analysis
What happened
Hugging Face has introduced Nunchaku Lite, which allows for 4-bit diffusion inference in Diffusers without the need for custom pipelines. This implementation reportedly achieves a 30% speed increase and reduces VRAM usage by 50%. The NVFP4 checkpoints can generate 1024x1024 images in approximately 1.7 seconds on an RTX 5090.
Key evidence
- Loading a modern text-to-image model in BF16 precision typically requires 20-30 GB of VRAM, limiting accessibility for consumer GPUs.
- Nunchaku Lite enables the use of 4-bit weights and activations, which reduces memory usage while speeding up the denoising loop.
- Benchmarks show that Nunchaku Lite NVFP4 can complete the full pipeline in 2.27 seconds with a peak VRAM usage of 20.6 GB, compared to 3.00 seconds and 31.1 GB for the BF16 baseline.
Why it matters
The introduction of Nunchaku Lite addresses the significant VRAM requirements of large diffusion models, making them more accessible to users with consumer-grade GPUs. By reducing both memory usage and inference time, it allows for more efficient image generation workflows. This could lead to broader adoption of diffusion models in various applications, including art generation and content creation.
Source Excerpt
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Hugging Face
See more →
From Hugging Face to Amazon SageMaker Studio in one click
Hugging Face has launched a deep-link integration with Amazon SageMaker Studio, allowing developers to seamlessly transition from model discovery to deployment with a single click. This integration streamlines the process by pre-configuring permissions and providing GPU quota visibility, significantly reducing the time from model selection to experimentation.




