跳到正文
北京时间
原文
Hacker News 热门(buzzing.cc 中文翻译)· Anon84·· 2026-08-04精选AI 评分76

AirLLM 实现单块 4GB GPU 运行 70B 模型推理

使用单块 4GB GPU 进行 AirLLM 70B 推理

AI 导读

AirLLM 项目支持在单块 4GB 显存 GPU 上运行 70B 参数大模型推理,无需多卡或大规模显存配置。该项目已开源,相关讨论在 Hacker News 上获得 103 点热度,引发社区关注。

推荐理由

将 70B 模型推理压缩到 4GB 显存,让不带专用显卡的开发者也能本地测试大模型,但实际推理速度和质量仍需根据用例评估。

正文 · AI 翻译

airllm_logo

AirLLM大幅降低推理内存占用,让 70B 大语言模型在单张 4GB GPU 显卡上运行——无需量化、知识蒸馏或剪枝。你甚至可以运行405B Llama 3.1在8GB, DeepSeek-V3(671B)在约 12GB,以及Kimi K3(2.8T)——迄今为止发布的最大开源模型——在不到 4GB上运行,因为稀疏 MoE 模型一次流式加载一个专家,而非整个层。

更新

[2026/07] Kimi K3(2.8T) 支持:这个最大的开源模型在单张卡上运行,仅占用 3.72GB 显存,在一张 RTX 6000 Ada 上完成了端到端实测。按专家流式加载,只加载 token 实际路由到的那些专家。K3 自身带来了三项要求:pip install compressed-tensors flash-attn(其模型代码强制要求 flash attention,无论你请求什么),一个 CUDA 12 构建的 torch,因为目前还没有针对 CUDA 13 的预编译 flash-attn wheel,以及 transformers 4.56.x,因为其远程代码无法在 5.x 上加载。

[2026/06] v3.0:支持 FP8 模型以及最新模型。可在约 12GB 显存上运行 DeepSeek-V3(671B),并在约 3GB 显存上运行 Qwen3-235B,此外还支持 Qwen3、Llama 3.x/4、DeepSeek V2/V3、Phi-4、Gemma 等更多模型——全部通过单一的 AutoModel 实现。

[2024/08/20] v2.11.0:支持 Qwen2.5

[2024/08/18] v2.10.1 支持 CPU 推理。支持非分片模型。感谢 @NavodPeiris 的出色工作!

[2024/07/30] 支持 Llama3.1 405B(示例 notebook)。支持 8bit/4bit 量化。

[2024/04/20] AirLLM 已原生支持 Llama3。在 4GB 单张 GPU 上运行 Llama3 70B。

[2023/12/25] v2.8.2:支持 MacOS 运行 70B 大语言模型。

[2023/12/20] v2.7:支持 AirLLMMixtral。

[2023/12/20] v2.6:新增 AutoModel,可自动检测模型类型,无需提供模型类即可初始化模型。

[2023/12/18] v2.5:新增预取功能,使模型加载与计算重叠进行。速度提升 10%。

[2023/12/03] 新增对 ChatGLM、QWen、Baichuan、Mistral、InternLM 的支持!

[2023/12/02] 新增对 safetensors 的支持。现已支持 open llm leaderboard 上排名前 10 的全部模型。

[2023/12/01] airllm 2.0。支持压缩:运行速度提升 3 倍!

[2023/11/20] airllm 初始版本!

Star 历史

Star History Chart

快速开始

1. 安装软件包

首先,安装 airllm pip 包。

pip install airllm

2. 推理

然后,初始化 AirLLMLlama2,传入所使用模型的 huggingface 仓库 ID 或本地路径,即可像普通 transformer 模型一样进行推理。

(在初始化 AirLLMLlama2 时,你还可以通过 layer_shards_saving_path 指定保存拆分后分层模型的路径。

from airllm import AutoModel

MAX_LENGTH = 128
# just pass a hugging face repo id — works with almost any popular model:
model = AutoModel.from_pretrained("Qwen/Qwen3-32B")

# go bigger with the exact same one line:
#model = AutoModel.from_pretrained("Qwen/Qwen3-235B-A22B")     # 235B, runs in ~3GB
#model = AutoModel.from_pretrained("deepseek-ai/DeepSeek-V3")  # 671B, runs in ~12GB

# or use a model's local path...
#model = AutoModel.from_pretrained("/home/ubuntu/.cache/huggingface/hub/models--Qwen--Qwen3-32B/snapshots/...")

input_text = [
        'What is the capital of United States?',
        #'I like',
    ]

input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=False)
           
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=20,
    use_cache=True,
    return_dict_in_generate=True)

output = model.tokenizer.decode(generation_output.sequences[0])

print(output)

注意:在推理过程中,原始模型会首先被分解并按层保存。请确保 huggingface 缓存目录中有足够的磁盘空间。

模型压缩——推理速度提升 3 倍!

我们刚刚加入了基于分块量化的模型压缩。它可以进一步加快推理速度,最高可达3 倍,而精度损失几乎可以忽略!(更多性能评估以及我们为何使用分块量化,请见这篇论文)

speed_improvement

如何启用模型压缩加速:

  • 第 1 步:确保你已安装bitsandbytes通过以下方式安装pip install -U bitsandbytes
  • 第 2 步。确保 airllm 版本高于 2.0.0:pip install -U airllm
  • 第 3 步。初始化模型时,传入参数 compression('4bit' 或 '8bit'):
model = AutoModel.from_pretrained("garage-bAInd/Platypus2-70B-instruct",
                     compression='4bit' # specify '8bit' for 8-bit block-wise quantization 
                    )

模型压缩与量化之间有什么区别?

量化通常需要同时对权重和激活值进行量化,才能真正实现加速。这使得保持精度、避免各种输入中离群值的影响变得更加困难。

在我们的场景中,瓶颈主要在于磁盘加载,因此我们只需要缩小模型加载的体积。所以,我们只需对权重部分进行量化,这样更容易保证精度。

配置

初始化模型时,我们支持以下配置:

  • compression:支持的选项:4bit、8bit,分别用于 4 位或 8 位分块量化,或默认 None 表示不压缩
  • profiling_mode:支持的选项:True 表示输出耗时,或默认 False
  • layer_shards_saving_path:可选,用于保存拆分后模型的另一个路径
  • hf_token:如果下载需要授权的模型,例如 meta-llama/Llama-2-7b-hf,可以在此处提供 huggingface token
  • prefetching:通过预取来重叠模型加载与计算。默认开启。目前只有 AirLLMLlama2 支持此功能。
  • delete_original:如果你的磁盘空间不太充裕,可以将 delete_original 设为 true,以删除原始下载的 hugging face 模型,只保留转换后的模型,从而节省一半磁盘空间。

MacOS

只需安装 airllm,然后像在 linux 上一样运行代码即可。更多内容请参见 Quick Start。

  • 请确保你已安装 mlx 和 torch
  • 你可能需要安装 python native,更多内容请参见 此处
  • 仅支持 Apple silicon

示例 [python notebook](https://github.com/lyogavin/airllm/blob/main/air_llm/examples/run_on_macos.ipynb)

示例 Python Notebook

示例 colab 见此处:

Open In Colab

其他模型示例(ChatGLM、QWen、Baichuan、Mistral 等):

  • ChatGLM:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("THUDM/chatglm3-6b-base")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH, 
    padding=True)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache= True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
  • QWen:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("Qwen/Qwen-7B")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache=True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])
  • Baichuan、InternLM、Mistral 等:
from airllm import AutoModel
MAX_LENGTH = 128
model = AutoModel.from_pretrained("baichuan-inc/Baichuan2-7B-Base")
#model = AutoModel.from_pretrained("internlm/internlm-20b")
#model = AutoModel.from_pretrained("mistralai/Mistral-7B-Instruct-v0.1")
input_text = ['What is the capital of China?',]
input_tokens = model.tokenizer(input_text,
    return_tensors="pt", 
    return_attention_mask=False, 
    truncation=True, 
    max_length=MAX_LENGTH)
generation_output = model.generate(
    input_tokens['input_ids'].cuda(), 
    max_new_tokens=5,
    use_cache=True,
    return_dict_in_generate=True)
model.tokenizer.decode(generation_output.sequences[0])

如需请求支持其他模型:点此

支持的模型

AirLLM 开箱即用,几乎支持所有主流开源 LLM——只需将其 Hugging Face ID 传给 AutoModel.from_pretrained(...) 即可。这涵盖了所有主要模型家族:

Llama(2 / 3 / 3.1 / 3.3 / 4)· Qwen(1 / 2 / 2.5 / 3,包括 MoE 和 FP8)· DeepSeek(V2 / V3 / R1)· Mistral & Mixtral · Phi · Gemma · ChatGLM · Baichuan · InternLM · Yi——以及大多数新模型在发布当天即可支持。

小 GPU,大模型

诀窍在于:AirLLM 每次只在 GPU 上保留一个层,因此你所需的 VRAM 取决于模型的层大小——而非其总大小。这就是 671B 模型能装进一张爱好者显卡的原因:

模型 大小 GPU VRAM
Qwen3 / Mistral / Phi(约 8B) 8B ~1–2 GB
Qwen3-30B / Mixtral(MoE) 30–47B ~1–3 GB
Qwen3-235B(MoE) 235B ~3 GB
Llama 3.x 70B(全精度) 70B ~4 GB
Llama 3.1 405B 405B ~8 GB
DeepSeek-V3 671B ~12 GB

它们全都用同一行代码搞定——无需任何特殊设置。

致谢

大量代码基于 SimJeg 在 Kaggle 考试竞赛中的出色工作。特别感谢 SimJeg:

GitHub 账号 @SimJeg、Kaggle 上的代码、相关讨论。

常见问题

1. MetadataIncompleteBuffer

safetensors_rust.SafetensorError: Error while deserializing header: MetadataIncompleteBuffer

如果你遇到这个错误,最可能的原因是你的磁盘空间用尽了。拆分模型的过程非常消耗磁盘。参见这个。你可能需要扩展磁盘空间,清除 huggingface .cache 然后重新运行。

2. ValueError: max() arg is an empty sequence

很可能你正在用 Llama2 类加载 QWen 或 ChatGLM 模型。请尝试以下方法:

对于 QWen 模型:

from airllm import AutoModel #<----- instead of AirLLMLlama2
AutoModel.from_pretrained(...)

对于 ChatGLM 模型:

from airllm import AutoModel #<----- instead of AirLLMLlama2
AutoModel.from_pretrained(...)

3. 401 Client Error....Repo model ... is gated.

有些模型是门控模型,需要 huggingface api token。你可以提供 hf_token:

model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-hf", #hf_token='HF_API_TOKEN')

4. ValueError: Asking to pad but the tokenizer does not have a padding token.

有些模型的 tokenizer 没有 padding token,因此你可以设置一个 padding token,或者干脆关闭 padding 配置:

input_tokens = model.tokenizer(input_text,
   return_tensors="pt", 
   return_attention_mask=False, 
   truncation=True, 
   max_length=MAX_LENGTH, 
   padding=False  #<-----------   turn off padding 
)

引用 AirLLM

如果你在研究中觉得 AirLLM 有用并希望引用它,请使用以下 BibTex 条目:

@software{airllm2023,
  author = {Gavin Li},
  title = {AirLLM: scaling large language models on low-end commodity computers},
  url = {https://github.com/lyogavin/airllm/},
  version = {0.0},
  year = {2023},
}

贡献

欢迎贡献、想法和讨论!

来源:Hacker News 热门(buzzing.cc 中文翻译) · github.com