跳到正文
北京时间
原文
Hugging Face:Blog·· 2025-05-21精选AI 评分60

Hugging Face Diffusers 量化后端对比:Flux-dev 内存与速度实测

Exploring Quantization Backends in Diffusers

AI 导读

Hugging Face 发布 Diffusers 量化后端指南,对比 bitsandbytes、torchao、Quanto、GGUF 和 FP8 Layerwise Casting 在 FLUX.1-dev 上把 BF16 加载内存约 31.447 GB 压缩到约 10.6 至 23.7 GB 的表现。

推荐理由

原文实测了 Diffusers 各量化后端在 Flux 上的内存和推理速度数据,读者可以按需求选后端并直接复用配置代码。

正文 · AI 翻译

Image 1: Hugging Face's logoHugging Face


返回文章

探索 Diffusers 中的量化后端

发布于 2025年5月21日

在 GitHub 上更新

- [x] 点赞 45

  • Image 2
  • Image 3
  • Image 4
  • Image 5
  • Image 6
  • Image 7
  • +39

Image 8: Derek Liu's avatar

Derek Liu derekl35 关注

Image 9: Marc Sun's avatar

Marc Sun marcsun13 关注

Image 10: Sayak Paul's avatar

Sayak Paul sayakpaul 关注

像 Flux(一种基于流的文本到图像生成模型)这样的大型扩散模型可以生成令人惊叹的图像,但其体积可能成为障碍,需要大量的内存和计算资源。量化提供了一种强大的解决方案,可以缩小这些模型,使其更易于使用,而不会大幅影响性能。但最大的问题始终是:你真的能分辨出最终图像的差异吗? 在我们深入探讨 Hugging Face Diffusers 中各种量化后端的技术细节之前,何不测试一下你自己的感知能力?

找出量化模型

我们创建了一个设置,你可以提供提示词,然后我们使用原始的高精度模型(例如 BF16 的 Flux-dev)和几个量化版本(BnB 4-bit、BnB 8-bit)生成结果。生成的图像会呈现给你,你的挑战是识别哪些来自量化模型。

在这里或下方试试吧!

通常,尤其是 8-bit 量化,差异很细微,不仔细检查可能察觉不到。像 4-bit 或更低这样的更激进的量化可能更明显,但结果仍然可以很好,尤其是考虑到巨大的内存节省。不过 NF4 通常能提供最佳的权衡。

现在,让我们深入探讨。

Diffusers 中的量化后端

基于我们之前的文章“使用 Quanto 和 Diffusers 实现内存高效的扩散 Transformer”,本文探讨了直接集成到 Hugging Face Diffusers 中的各种量化后端。我们将研究 bitsandbytes、GGUF、torchao、Quanto 和原生 FP8 支持如何使大型强大的模型更易于使用,并以 Flux 为例进行演示。

在深入量化后端之前,让我们先介绍 FluxPipeline(使用 black-forest-labs/FLUX.1-dev 检查点)及其组件,我们将对其进行量化。以 BF16 精度加载完整的 FLUX.1-dev 模型大约需要 31.447 GB 内存。主要组件包括:

  • 文本编码器(CLIP 和 T5):

    • 功能:处理输入文本提示。FLUX-dev 使用 CLIP 进行初步理解,并使用更大的 T5 进行细致理解和更好的文本渲染。
    • 内存:T5 - 9.52 GB;CLIP - 246 MB(BF16 格式)
  • Transformer(主模型 - MMDiT):

    • 功能:核心生成部分(多模态扩散 Transformer)。从文本嵌入在潜在空间中生成图像。
    • 内存:23.8 GB(BF16 格式)
  • 变分自编码器(VAE):

    • 功能:在像素空间和潜在空间之间转换图像。将生成的潜在表示解码为基于像素的图像。
    • 内存:168 MB(BF16 格式)
  • 量化重点:示例将主要关注 transformer 和 text_encoder_2(T5),以实现最大的内存节省。

prompts = [
    "Baroque style, a lavish palace interior with ornate gilded ceilings, intricate tapestries, and dramatic lighting over a grand staircase.",
    "Futurist style, a dynamic spaceport with sleek silver starships docked at angular platforms, surrounded by distant planets and glowing energy lines.",
    "Noir style, a shadowy alleyway with flickering street lamps and a solitary trench-coated figure, framed by rain-soaked cobblestones and darkened storefronts.",
]

bitsandbytes (BnB)

bitsandbytes 是一个流行且用户友好的库,用于 8 位和 4 位量化,广泛用于 LLM 和 QLoRA 微调。我们也可以将其用于基于 transformer 的扩散和流模型。

BF16 BnB 4-bit BnB 8-bit 使用 BF16(左)、BnB 4-bit(中)和 BnB 8-bit(右)量化的 Flux-dev 模型输出视觉比较。(点击图像放大)

精度 加载后内存 峰值内存 推理时间
BF16 ~31.447 GB 36.166 GB 12 秒
4-bit 12.584 GB 17.281 GB 12 秒
8-bit 19.273 GB 24.432 GB 27 秒

所有基准测试均在 1x NVIDIA H100 80GB GPU 上执行

示例(Flux-dev 使用 BnB 4-bit):

import torch
from diffusers import FluxPipeline
from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
from diffusers.quantizers import PipelineQuantizationConfig
from transformers import BitsAndBytesConfig as TransformersBitsAndBytesConfig

model_id = "black-forest-labs/FLUX.1-dev"

pipeline_quant_config = PipelineQuantizationConfig(
    quant_mapping={
        "transformer": DiffusersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
        "text_encoder_2": TransformersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
    }
)

pipe = FluxPipeline.from_pretrained(
    model_id,
    quantization_config=pipeline_quant_config,
    torch_dtype=torch.bfloat16
)
pipe.to("cuda")

prompt = "Baroque style, a lavish palace interior with ornate gilded ceilings, intricate tapestries, and dramatic lighting over a grand staircase."
pipe_kwargs = {
    "prompt": prompt,
    "height": 1024,
    "width": 1024,
    "guidance_scale": 3.5,
    "num_inference_steps": 50,
    "max_sequence_length": 512,
}

print(f"Pipeline memory usage: {torch.cuda.max_memory_reserved() / 1024**3:.3f} GB")

image = pipe(
    **pipe_kwargs, generator=torch.manual_seed(0),
).images[0]

print(f"Pipeline memory usage: {torch.cuda.max_memory_reserved() / 1024**3:.3f} GB")

image.save("flux-dev_bnb_4bit.png")

注意:当使用 PipelineQuantizationConfig 与 bitsandbytes 时,您需要分别从 diffusers 导入 DiffusersBitsAndBytesConfig,从 transformers 导入 TransformersBitsAndBytesConfig。这是因为这些组件来自不同的库。如果您更喜欢更简单的设置,而不需要管理这些不同的导入,您可以使用管道级量化的替代方法,此方法的示例在 Diffusers 文档中的管道级量化。

更多信息请查看 bitsandbytes 文档。

torchao

torchao 是一个 PyTorch 原生的库,用于架构优化,提供量化、稀疏性和自定义数据类型,设计用于与 torch.compile 和 FSDP 兼容。Diffusers 支持广泛的 torchao 的奇异数据类型,实现对模型优化的细粒度控制。

int4_weight_only int8_weight_only float8_weight_only 使用 torchao int4_weight_only(左)、int8_weight_only(中)和 float8_weight_only(右)量化的 Flux-dev 模型输出视觉比较。(点击图像放大)

torchao 精度 加载后内存 峰值内存 推理时间
int4_weight_only 10.635 GB 14.654 GB 109 秒
int8_weight_only 17.020 GB 21.482 GB 15 秒
float8_weight_only 17.016 GB 21.488 GB 15 秒

示例(Flux-dev 使用 torchao INT8 weight-only):

@@
- from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
+ from diffusers import TorchAoConfig as DiffusersTorchAoConfig

- from transformers import BitsAndBytesConfig as TransformersBitsAndBytesConfig
+ from transformers import TorchAoConfig as TransformersTorchAoConfig
@@
pipeline_quant_config = PipelineQuantizationConfig(
    quant_mapping={
-         "transformer": DiffusersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
-         "text_encoder_2": TransformersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
+         "transformer": DiffusersTorchAoConfig("int8_weight_only"),
+         "text_encoder_2": TransformersTorchAoConfig("int8_weight_only"),
    }
)

示例(Flux-dev 使用 torchao INT4 weight-only):

@@
- from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
+ from diffusers import TorchAoConfig as DiffusersTorchAoConfig

- from transformers import BitsAndBytesConfig as TransformersBitsAndBytesConfig
+ from transformers import TorchAoConfig as TransformersTorchAoConfig
@@
pipeline_quant_config = PipelineQuantizationConfig(
    quant_mapping={
-         "transformer": DiffusersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
-         "text_encoder_2": TransformersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
+         "transformer": DiffusersTorchAoConfig("int4_weight_only"),
+         "text_encoder_2": TransformersTorchAoConfig("int4_weight_only"),
    }
)

pipe = FluxPipeline.from_pretrained(
    model_id,
    quantization_config=pipeline_quant_config,
    torch_dtype=torch.bfloat16,
+    device_map="balanced"
)
- pipe.to("cuda")

更多信息请查看 torchao 文档。

Quanto

Quanto 是一个通过 optimum 库与 Hugging Face 生态系统集成的量化库。

INT4 INT8 FP8 使用 Quanto INT4(左)、INT8(中)和 FP8(右)量化的 Flux-dev 模型输出视觉比较。(点击图像放大)

quanto 精度 加载后内存 峰值内存 推理时间
INT4 12.254 GB 16.139 GB 109 秒
INT8 17.330 GB 21.814 GB 15 秒
FP8 16.395 GB 20.898 GB 16 秒

示例(Flux-dev 使用 quanto INT8 weight-only):

@@
- from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
+ from diffusers import QuantoConfig as DiffusersQuantoConfig

- from transformers import BitsAndBytesConfig as TransformersBitsAndBytesConfig
+ from transformers import QuantoConfig as TransformersQuantoConfig
@@
pipeline_quant_config = PipelineQuantizationConfig(
    quant_mapping={
-         "transformer": DiffusersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
-         "text_encoder_2": TransformersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
+         "transformer": DiffusersQuantoConfig(weights_dtype="int8"),
+         "text_encoder_2": TransformersQuantoConfig(weights_dtype="int8"),
    }
)

注意:在撰写本文时,若需 Quanto 的 float8 支持,您需要 optimum-quanto<0.2.5 并直接使用 quanto。我们正在努力修复此问题。

示例(使用 quanto FP8 仅权重的 Flux-dev)

import torch
from diffusers import AutoModel, FluxPipeline
from transformers import T5EncoderModel
from optimum.quanto import freeze, qfloat8, quantize

model_id = "black-forest-labs/FLUX.1-dev"

text_encoder_2 = T5EncoderModel.from_pretrained(
    model_id,
    subfolder="text_encoder_2",
    torch_dtype=torch.bfloat16,
)

quantize(text_encoder_2, weights=qfloat8)
freeze(text_encoder_2)

transformer = AutoModel.from_pretrained(
      model_id,
      subfolder="transformer",
      torch_dtype=torch.bfloat16,
)

quantize(transformer, weights=qfloat8)
freeze(transformer)

pipe = FluxPipeline.from_pretrained(
    model_id,
    transformer=transformer,
    text_encoder_2=text_encoder_2,
    torch_dtype=torch.bfloat16
).to("cuda")

更多信息请查阅 Quanto 文档。

GGUF

GGUF 是 llama.cpp 社区中用于存储量化模型的流行文件格式。

Q2_k Q4_1 Q8_0 使用 GGUF Q2_k(左)、Q4_1(中)和 Q8_0(右)量化的 Flux-dev 模型输出视觉对比。(点击图片可放大)

GGUF 精度 加载后内存 峰值内存 推理时间
Q2_k 13.264 GB 17.752 GB 26 秒
Q4_1 16.838 GB 21.326 GB 23 秒
Q8_0 21.502 GB 25.973 GB 15 秒

示例(使用 GGUF Q4_1 的 Flux-dev)

import torch
from diffusers import FluxPipeline, FluxTransformer2DModel, GGUFQuantizationConfig

model_id = "black-forest-labs/FLUX.1-dev"

# Path to a pre-quantized GGUF file
ckpt_path = "https://huggingface.co/city96/FLUX.1-dev-gguf/resolve/main/flux1-dev-Q4_1.gguf"

transformer = FluxTransformer2DModel.from_single_file(
    ckpt_path,
    quantization_config=GGUFQuantizationConfig(compute_dtype=torch.bfloat16),
    torch_dtype=torch.bfloat16,
)

pipe = FluxPipeline.from_pretrained(
    model_id,
    transformer=transformer,
    torch_dtype=torch.bfloat16,
)
pipe.to("cuda")

更多信息请查阅 GGUF 文档。

FP8 逐层转换(enable_layerwise_casting)

FP8 逐层转换是一种内存优化技术。它通过将模型权重以紧凑的 FP8(8 位浮点)格式存储来工作,这种格式使用的内存大约是标准 FP16 或 BF16 精度的一半。就在某一层执行计算之前,其权重会动态转换为更高的计算精度(如 FP16/BF16)。紧接着,权重又被转换回 FP8 以实现高效存储。这种方法之所以有效,是因为核心计算保持了高精度,并且对量化特别敏感的层(如归一化)通常会被跳过。此技术还可以与 组卸载 结合使用,以进一步节省内存。

FP8 (e4m3) 使用 FP8 逐层转换(e4m3)量化的 Flux-dev 模型视觉输出。

精度 加载后内存 峰值内存 推理时间
FP8 (e4m3) 23.682 GB 28.451 GB 13 秒
import torch
from diffusers import AutoModel, FluxPipeline

model_id = "black-forest-labs/FLUX.1-dev"

transformer = AutoModel.from_pretrained(
    model_id,
    subfolder="transformer",
    torch_dtype=torch.bfloat16
)
transformer.enable_layerwise_casting(storage_dtype=torch.float8_e4m3fn, compute_dtype=torch.bfloat16)

pipe = FluxPipeline.from_pretrained(model_id, transformer=transformer, torch_dtype=torch.bfloat16)
pipe.to("cuda")

更多信息请查阅 逐层转换文档。

结合更多内存优化和 torch.compile

大多数这些量化后端都可以与 Diffusers 中提供的内存优化技术结合使用。让我们探索 CPU 卸载、组卸载和 torch.compile。您可以在 Diffusers 文档 中了解更多关于这些技术的信息。

注意:在撰写本文时,如果从源代码安装 bnb 并使用 pytorch nightly 或 fullgraph=False,bnb + torch.compile 也能工作。

示例(使用 BnB 4 位 + enable_model_cpu_offload 的 Flux-dev):

import torch
from diffusers import FluxPipeline
from diffusers import BitsAndBytesConfig as DiffusersBitsAndBytesConfig
from diffusers.quantizers import PipelineQuantizationConfig
from transformers import BitsAndBytesConfig as TransformersBitsAndBytesConfig

model_id = "black-forest-labs/FLUX.1-dev"

pipeline_quant_config = PipelineQuantizationConfig(
    quant_mapping={
        "transformer": DiffusersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
        "text_encoder_2": TransformersBitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16),
    }
)

pipe = FluxPipeline.from_pretrained(
    model_id,
    quantization_config=pipeline_quant_config,
    torch_dtype=torch.bfloat16
)
- pipe.to("cuda")
+ pipe.enable_model_cpu_offload()

模型 CPU 卸载(enable_model_cpu_offload):此方法在推理管道期间将整个模型组件(如 UNet、文本编码器或 VAE)在 CPU 和 GPU 之间移动。它提供了显著的 VRAM 节省,并且通常比更细粒度的卸载更快,因为它涉及更少、更大的数据传输。

bnb + enable_model_cpu_offload:

精度 加载后内存 峰值内存 推理时间
4 位 12.383 GB 12.383 GB 17 秒
8 位 19.182 GB 23.428 GB 27 秒

示例(使用 fp8 逐层转换 + 组卸载的 Flux-dev):

import torch
from diffusers import FluxPipeline, AutoModel

model_id = "black-forest-labs/FLUX.1-dev"

transformer = AutoModel.from_pretrained(
    model_id,
    subfolder="transformer",
    torch_dtype=torch.bfloat16,
    # device_map="cuda"
)
transformer.enable_layerwise_casting(storage_dtype=torch.float8_e4m3fn, compute_dtype=torch.bfloat16)
+ transformer.enable_group_offload(onload_device=torch.device("cuda"), offload_device=torch.device("cpu"), offload_type="leaf_level", use_stream=True)

pipe = FluxPipeline.from_pretrained(model_id, transformer=transformer, torch_dtype=torch.bfloat16)
- pipe.to("cuda")

组卸载(enable_group_offload 用于 diffusers 组件或 apply_group_offloading 用于通用 torch.nn.Module):它将内部模型层的组(如 torch.nn.ModuleList 或 torch.nn.Sequential 实例)移动到 CPU。这种方法通常比完整模型卸载更节省内存,并且比顺序卸载更快。

FP8 逐层转换 + 组卸载:

精度 加载后内存 峰值内存 推理时间
FP8 (e4m3) 9.264 GB 14.232 GB 58 秒

示例(使用 torchao 4 位 + torch.compile 的 Flux-dev):

import torch
from diffusers import FluxPipeline
from diffusers import TorchAoConfig as DiffusersTorchAoConfig
from diffusers.quantizers import PipelineQuantizationConfig
from transformers import TorchAoConfig as TransformersTorchAoConfig

from torchao.quantization import Float8WeightOnlyConfig

model_id = "black-forest-labs/FLUX.1-dev"
dtype = torch.bfloat16

pipeline_quant_config = PipelineQuantizationConfig(
    quant_mapping={
        "transformer":DiffusersTorchAoConfig("int4_weight_only"),
        "text_encoder_2": TransformersTorchAoConfig("int4_weight_only"),
    }
)

pipe = FluxPipeline.from_pretrained(
    model_id,
    quantization_config=pipeline_quant_config,
    torch_dtype=torch.bfloat16,
    device_map="balanced"
)

+ pipe.transformer = torch.compile(pipe.transformer, mode="max-autotune", fullgraph=True)

注意:torch.compile 可能会引入细微的数值差异,导致图像输出变化

torch.compile:另一种互补的方法是使用 PyTorch 2.x 的 torch.compile() 功能来加速模型执行。编译模型不会直接降低内存,但可以显著加快推理速度。PyTorch 2.0 的 compile(Torch Dynamo)通过提前追踪和优化模型图来工作。

torchao + torch.compile:

torchao 精度 加载后内存 峰值内存 推理时间 编译时间
int4_weight_only 10.635 GB 15.238 GB 6 秒 约 285 秒
int8_weight_only 17.020 GB 22.473 GB 8 秒 约 851 秒
float8_weight_only 17.016 GB 22.115 GB 8 秒 约 545 秒

在此处查看一些基准测试结果:

可直接使用的量化检查点

你可以在我们的 Hugging Face 集合中找到这篇博客文章中的 bitsandbytes 和 torchao 量化模型:集合链接。

结论

以下是选择量化后端的快速指南:

  • 最轻松节省内存(NVIDIA):从 bitsandbytes 4/8 位开始。这也可以与 torch.compile() 结合使用以加快推理速度。
  • 优先考虑推理速度:torchao、GGUF 和 bitsandbytes 都可以与 torch.compile() 一起使用,以可能提升推理速度。
  • 为了硬件灵活性(CPU/MPS)、FP8 精度:Quanto 可能是一个不错的选择。
  • 简单性(Hopper/Ada):探索 FP8 逐层转换(enable_layerwise_casting)。
  • 使用现有 GGUF 模型:使用 GGUF 加载(from_single_file)。
  • 对量化训练感到好奇?请期待后续关于该主题的博客文章!更新(2025 年 6 月 19 日):它在这里!

量化显著降低了使用大型扩散模型的门槛。尝试这些后端,为你的需求找到内存、速度和质量的最佳平衡。

致谢:感谢 Chunte 为本文提供缩略图。

本文提到的模型 1

Image 11 #### black-forest-labs/FLUX.1-dev 文本到图像 • 12B•更新于 2025 年 6 月 27 日• 713k• 15.3k

本文提到的 Space 1

已暂停 Agents 9 #### Flux 量化还是原始?🔎 9 根据提示生成并比较量化图像

本文提到的集合 1

Flux 量化检查点集合 此集合汇总了我们在本博客文章中使用的量化 flux 检查点:https://huggingface.co/blog/diffusers-quantization•5 个项目•更新于 2025 年 11 月 26 日• 2

我们博客中的更多文章

Image 12 指南 diffusers 量化 ## 在消费级硬件上(LoRA)微调 FLUX.1-dev * Image 13 * Image 14 * Image 15 * Image 16 * +1 107 2025 年 6 月 19 日 derekl35 等

Image 17 指南 diffusers 量化 ## 将 Nunchaku 4 位扩散推理引入 Diffusers * Image 18 * Image 19 69 2026 年 7 月 23 日 rootonchair 等

社区

Image 20

tolgacangoz

2025 年 10 月 25 日

•

此评论已被隐藏(标记为已解决)

编辑预览

通过拖入文本输入框、粘贴或点击此处来上传图像、音频和视频。

点击或粘贴此处以上传图像

评论 ·注册或登录以发表评论

- [x] 点赞 45

  • Image 21
  • Image 22
  • Image 23
  • Image 24
  • Image 25
  • Image 26
  • Image 27
  • Image 28
  • Image 29
  • Image 30
  • Image 31
  • Image 32
  • +33

本文提到的模型 1

Image 33 #### black-forest-labs/FLUX.1-dev 文本到图像 • 12B•更新于 2025 年 6 月 27 日• 713k• 15.3k

本文提到的 Space 1

已暂停 Agents 9 #### Flux 量化还是原始?🔎 9 根据提示生成并比较量化图像

本文提到的集合 1

Flux 量化检查点合集 该合集汇总了我们在本篇博客文章中使用的量化 flux 检查点:https://huggingface.co/blog/diffusers-quantization•5 个项目•更新于 2025年11月26日• 2

系统主题

公司

服务条款隐私关于招聘

网站

模型数据集Spaces定价文档

来源:Hugging Face:Blog · huggingface.co