跳到正文
北京时间
原文
Hugging Face:Blog·· 2025-06-26精选AI 评分81

Gemma 3n 开源生态全面可用,发布 E2B 和 E4B 两种尺寸

Gemma 3n fully available in the open-source ecosystem!

AI 导读

Gemma 3n 在 transformers、timm、MLX、llama.cpp、Transformers.js、Ollama 和 Google AI Edge 等开源库中全面可用,发布 gemma-3n-E2B 和 gemma-3n-E4B 两个尺寸,各有 base 与 instruct 变体。

推荐理由

原文给出 E2B/E4B 的显存占用、架构细节和各库用法与微调入口,读者可以直接照着跑起来。

正文 · AI 翻译

Gemma 3n 在 Google I/O 上作为预览版发布。设备端社区非常兴奋,因为这是一个从底层设计就旨在本地运行于你的硬件上的模型。更重要的是,它原生支持多模态,可处理图像、文本、音频和视频输入 🤯

今天,Gemma 3n 终于可以在最常用的开源库中使用了。这包括 transformers & timm、MLX、llama.cpp(文本输入)、transformers.js、ollama、Google AI Edge 等。

本文通过实用代码片段快速演示如何在这些库中使用该模型,以及如何轻松地针对其他领域对其进行微调。

今日发布的模型

这里是 Gemma 3n 发布合集

今天发布了两种模型规模,每种都有两个变体(基础版和指令版)。模型名称遵循非标准命名法:它们被称为 gemma-3n-E2B 和 gemma-3n-E4B。参数数量前的 E 代表 Effective。它们的实际参数数量分别为 5B 和 8B,但得益于内存效率的改进,它们仅需 2B 和 4B 的 VRAM(GPU 内存)。

因此,这些模型在硬件支持方面表现得像 2B 和 4B,但在质量上却超越了 2B/4B。E2B 模型仅需 2GB 的 GPU RAM 即可运行,而 E4B 只需 3GB 的 GPU RAM 即可运行。

规模 基础版 指令版
2B google/gemma-3n-e2b google/gemma-3n-e2b-it
4B google/gemma-3n-e4b google/gemma-3n-e4b-it

模型详情

除了语言解码器外,Gemma 3n 还使用了一个音频编码器和一个视觉编码器。下面我们重点介绍它们的主要特性,并描述它们是如何被添加到 transformers 和 timm 中的,因为它们是其他实现的参考。

  • Vision Encoder (MobileNet-V5). Gemma 3n uses a new version of MobileNet: MobileNet-v5-300, which has been added to the new version of timm released today.
    • 具有 3 亿参数。
    • 支持 256x256、512x512 和 768x768 分辨率。
    • 在 Google Pixel 上达到 60 FPS,性能超越 ViT Giant,同时使用的参数少 3 倍。
  • Audio Encoder:
    • 基于通用语音模型(USM)。
    • 以 160ms 块处理音频。
    • 支持语音转文本和翻译功能(例如,英语到西班牙语/法语)。
  • Gemma 3n 架构与语言模型。该架构本身已添加到今天发布的新版 transformers 中。此实现分支到 timm 进行图像编码,因此 MobileNet 架构有一个单一的参考实现。

架构亮点

  • MatFormer Architecture:
    • 类似 Matryoshka 嵌入的嵌套 Transformer 设计,允许提取各种层子集,就像它们是独立模型一样。
    • E2B 和 E4B 一起训练,将 E2B 配置为 E4B 的子模型。
    • 用户可以根据硬件特性和内存预算“混合搭配”层。
  • 逐层嵌入(PLE):通过将嵌入卸载到 CPU 来减少加速器内存使用。这就是为什么 E2B 模型虽然拥有 5B 实际参数,却占用与 2B 参数模型大致相同的 GPU 内存。
  • KV 缓存共享:加速音频和视频的长上下文处理,与 Gemma 3 4B 相比,预填充速度提高 2 倍。

性能与基准测试:

  • LMArena 得分:E4B 是首个得分超过 1300 的 10B 以下模型。
  • MMLU 分数:Gemma 3n 在各种规模(E4B、E2B 以及多种 Mix-n-Match 配置)下都展现出有竞争力的性能。
  • 多语言支持:支持 140 种语言的文本和 35 种语言的多模态交互。

演示空间

GIF of Hugging Face Space for Gemma 3n

对模型进行 vibe check 最简单的方法是使用该模型的专用 Hugging Face Space。你可以在这里尝试不同的提示词和不同的模态。

📱 Space

使用 transformers 进行推理

安装最新版本的 timm(用于视觉编码器)和 transformers 以运行推理,或者如果你想对其进行微调。

pip install -U -q timm
pip install -U -q transformers

使用 pipeline 进行推理

开始使用 Gemma 3n 最简单的方法是使用 transformers 中的 pipeline 抽象:

import torch
from transformers import pipeline

pipe = pipeline(
   "image-text-to-text",
   model="google/gemma-3n-E4B-it", # "google/gemma-3n-E4B-it"
   device="cuda",
   torch_dtype=torch.bfloat16
)

messages = [
   {
       "role": "user",
       "content": [
           {"type": "image", "url": "https://huggingface.co/datasets/ariG23498/demo-data/resolve/main/airplane.jpg"},
           {"type": "text", "text": "Describe this image"}
       ]
   }
]

output = pipe(text=messages, max_new_tokens=32)
print(output[0]["generated_text"][-1]["content"])

输出:

The image shows a futuristic, sleek aircraft soaring through the sky. It's designed with a distinctive, almost alien aesthetic, featuring a wide body and large

使用 transformers 进行详细推理

从 Hub 初始化模型和处理器,并编写 model_generation 函数来处理提示词并在模型上运行推理。

from transformers import AutoProcessor, AutoModelForImageTextToText
import torch

model_id = "google/gemma-3n-e4b-it" # google/gemma-3n-e2b-it
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id).to(device)

def model_generation(model, messages):
    inputs = processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=True,
        return_dict=True,
        return_tensors="pt",
    )
    input_len = inputs["input_ids"].shape[-1]

    inputs = inputs.to(model.device, dtype=model.dtype)

    with torch.inference_mode():
        generation = model.generate(**inputs, max_new_tokens=32, disable_compile=False)
        generation = generation[:, input_len:]

    decoded = processor.batch_decode(generation, skip_special_tokens=True)
    print(decoded[0])

由于该模型支持所有模态作为输入,这里简要说明如何通过 transformers 使用它们。

仅文本

# Text Only

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "What is the capital of France?"}
        ]
    }
]
model_generation(model, messages)

输出:

The capital of France is **Paris**. 

与音频交错

# Interleaved with Audio

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Transcribe the following speech segment in English:"},
            {"type": "audio", "audio": "https://huggingface.co/datasets/ariG23498/demo-data/resolve/main/speech.wav"},
        ]
    }
]
model_generation(model, messages)

输出:

Send a text to Mike. I'll be home late tomorrow.

与图像/视频交错

对视频的支持是通过图像帧集合来实现的

# Interleaved with Image

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/ariG23498/demo-data/resolve/main/airplane.jpg"},
            {"type": "text", "text": "Describe this image."}
        ]
    }
]
model_generation(model, messages)

输出:

The image shows a futuristic, sleek, white airplane against a backdrop of a clear blue sky transitioning into a cloudy, hazy landscape below. The airplane is tilted at

使用 MLX 进行推理

Gemma 3n 在所有 3 种模态上都提供 MLX 的 day 0 支持。请确保升级你的 mlx-vlm 安装。

pip install -u mlx-vlm

视觉入门:

python -m mlx_vlm.generate --model google/gemma-3n-E4B-it --max-tokens 100 --temperature 0.5 --prompt "Describe this image in detail." --image https://huggingface.co/datasets/ariG23498/demo-data/resolve/main/airplane.jpg

以及音频:

python -m mlx_vlm.generate --model google/gemma-3n-E4B-it --max-tokens 100 --temperature 0.0 --prompt "Transcribe the following speech segment in English:" --audio https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/audio-samples/jfk.wav

使用 llama.cpp 进行推理

除了 MLX 之外,Gemma 3n(仅文本)还可以开箱即用地与 llama.cpp 配合使用。请确保从源码安装 llama.cpp/Ollama。

在此查看 llama.cpp 的安装说明:https://github.com/ggml-org/llama.cpp/blob/master/docs/install.md

你可以这样运行:

llama-server -hf ggml-org/gemma-3n-E4B-it-GGUF:Q8_0

使用 Transformers.js 和 ONNXRuntime 进行推理

最后,我们还发布了 gemma-3n-E2B-it 模型变体的 ONNX 权重,从而支持在不同运行时和平台上灵活部署。对于 JavaScript 开发者,Gemma3n 已集成到 Transformers.js 中,并自 3.6.0 版本起可用。

有关如何使用这些库运行该模型的更多信息,请查看 模型卡 中的使用部分。

在免费的 Google Colab 中微调

鉴于该模型的规模,针对跨模态的特定下游任务对其进行微调非常方便。为了让你更轻松地微调该模型,我们创建了一个简单的 notebook,让你可以在免费的 Google Colab 上进行实验!

我们还提供了一个专用的 音频任务微调 notebook,这样你就可以轻松地将模型适配到你的语音数据集和基准测试中!

Hugging Face Gemma Recipes

随着本次发布,我们还推出了 Hugging Face Gemma Recipes 仓库。你可以在其中找到用于运行和微调这些模型的 notebooks 和 scripts。

我们非常希望你能使用 Gemma 系列模型,并为其添加更多 recipes!欢迎随时向该仓库提交 Issues 和创建 Pull Requests。

结论

我们一直很高兴能托管 Google 及其 Gemma 系列模型。我们希望社区能够齐心协力,充分利用这些模型。多模态、小尺寸且能力强大,这是一次很棒的模型发布!

如果你想更详细地讨论这些模型,请在这篇博客文章下方发起讨论。我们将非常乐意提供帮助!

非常感谢 Arthur、Cyril、Raushan、Lysandre 以及 Hugging Face 的所有人,他们负责了集成工作,并将其提供给社区!

来源:Hugging Face:Blog · huggingface.co