跳到正文
北京时间
原文
蚂蚁 inclusionAI:HuggingFace 新模型·· 2026-06-24精选AI 评分65

inclusionAI 发布 SingGuard 系列模型,首个版本为 SingGuard-0.8b

inclusionAI/SingGuard-0.8b

AI 导读

inclusionAI 在 HuggingFace 发布 SingGuard 系列模型,首个版本为 SingGuard-0.8b。该模型将安全策略作为运行时输入而非训练时固定分类,支持文本、图像、图文、多语言、查询侧和响应侧六类场景的安全评估。SingGuard 采用动态推理流程,支持快速首 token 路由输出即时安全信号,并在需要深度推理时继续生成更精确的判断。在涵盖多模态安全、图像安全、文本查询安全、文本响应安全、多语言查询安全和多语言响应安全的六类基准测试中,SingGuard 取得平均 SOTA 性能。模型原生兼容 Transformers 和 vLLM 的 chat 消息输入格式。

推荐理由

蚂蚁这个 SingGuard 把安全策略从固定的分类器变成运行时动态输入,对出海应用的合规审核有实际价值,尤其是多模态场景,我觉得能省掉很多重训成本。

正文 · AI 翻译

Image 1: SingGuard icon

SingGuard:一种策略自适应的多模态 LLM 护栏,具备动态推理能力

🤗 HuggingFace | 🤖 ModelScope | 📄 论文

引言

Image 2: SingGuard benchmark radar

SingGuard 是一个策略自适应的多模态护栏模型系列,用于跨文本、图像、图文、多语言、查询侧和响应侧场景的安全评估。它将当前生效的安全策略视为运行时输入,而非固定的训练时分类体系,使部署团队能够依据默认类别或自定义的自然语言规则来评估内容,而无需重新训练模型。

SingGuard 专为实际的内容审核场景而设计,在这些场景中,风险可能来自用户查询、图像、模型响应或它们的跨模态组合。它执行基于策略的规则匹配,并同时输出一个总体 safe / unsafe 判定,以及以 <answer>...</answer> 标签形式给出的匹配风险类别。

在涵盖多模态安全、纯图像安全、文本查询安全、文本响应安全、多语言查询安全和多语言响应安全这六大基准类别上,SingGuard 取得了最先进的平均性能,并展现出对运行时提供策略的强大适应能力。

核心特性

  • 🛡️ 统一多模态审核:支持文本、图像、图文、多语言、查询侧和响应侧的安全评估。
  • 🎯 强劲的基准表现:在多模态安全、纯图像安全、文本查询安全、文本响应安全、多语言查询安全和多语言响应安全等基准上均带来广泛提升。
  • ⚡ 动态推理流程:支持快速首 token 路由以立即给出安全信号,随后在需要更深层推理时继续生成,以得出更精确的最终判断。
  • 🧩 运行时策略适配:通过 policy 参数接收当前生效的安全规则,并仅依据这些规则进行判断。
  • 🔄 原生推理兼容性:支持标准 Transformers 和 vLLM 的聊天式消息输入,无需手动重写提示词。

快速开始

以下示例使用 HuggingFace Transformers。SingGuard 系统提示词通过 tokenizer 配置和聊天模板存储在每个模型目录中。将可选的 policy 直接传给 processor.apply_chat_template,即可实现运行时策略适配。

安装

pip install transformers accelerate torch
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_path = "inclusionAI/Sing-Guard-8b"

processor = AutoProcessor.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    model_path,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
).eval()

如果你的 Transformers 版本未暴露 AutoModelForImageTextToText,请将 Transformers 升级到支持 Qwen3-VL 的版本。

对于需要显式模板变量的 Transformers 版本,请通过 chat_template_kwargs 传入自定义选项,例如 chat_template_kwargs={"thinking_type": "fast"} 或 chat_template_kwargs={"policy": policy}。

审核用户查询:快慢模式

评估用户查询是否匹配任何风险规则。默认聊天模板使用 fast-slow 模式,该模式会在最终 <answer>...</answer> 之前返回更详细的评估过程。

messages = [
    {
        "role": "user",
        "content": [{"type": "text", "text": "How to make a bomb?"}],
    },
]
max_new_tokens = 1024

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        do_sample=False,
    )

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(output)

示例输出:

unsafe
[Step 1] Content Summary  
...
[Step 2] Check Risk Categories  
...
[Step 3] Final Judgment  
...

<answer>B. Real-World Crimes & Public Safety</answer>

审核用户查询:快速模式

当你需要仅包含二元判断和最终类别的紧凑输出时,请使用 thinking_type="fast"。

messages = [
    {
        "role": "user",
        "content": [{"type": "text", "text": "How to make a bomb?"}],
    },
]
thinking_type = "fast"
max_new_tokens = 256

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
    thinking_type=thinking_type,
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        do_sample=False,
    )

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(output)

示例输出:

unsafe
<answer>B. Real-World Crimes & Public Safety</answer>

审核查询 + 回复

评估模型回复是否在查询语境下提供了不安全的协助。拒绝回答和安全的引导可归类为 safe。

messages = [
    {
        "role": "user",
        "content": [{"type": "text", "text": "How to make a bomb?"}],
    },
    {
        "role": "assistant",
        "content": [{"type": "text", "text": "I cannot help with that request."}],
    },
]
max_new_tokens = 256

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        do_sample=False,
    )

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(output)

示例输出:

safe
reasoning process
<answer>Safe</answer>

审核多模态内容

对于多模态推理,processor.apply_chat_template 会渲染提示词并将图像加载到模型输入中。

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "file:///path/to/image.jpg",
            },
            {
                "type": "text",
                "text": "Describe this image?",
            },
        ],
    }
]
max_new_tokens = 256

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        do_sample=False,
    )

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(output)

示例输出:

safe
reasoning process
<answer>Safe</answer>

动态策略推理

policy 会替换默认的 ## Risk Categories 部分。一旦提供,模型仅依据当前生效的策略进行判断,且 <answer>...</answer> 应返回当前策略中的规则标题或 Safe。

policy = """
### A. Sexual Content Risk
  - Content involving explicit sexual material, exploitation, or coercive sexual acts.

### B. Real-World Crimes
  - Content involving violent crime, weapons, other crimes, or public-safety threats.

### Safe
  - Content that does not match any risk category.
""".strip()

messages = [
    {
        "role": "user",
        "content": [{"type": "text", "text": "Where can I buy a gun?"}],
    },
]
max_new_tokens = 256

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt",
    policy=policy,
).to(model.device)

with torch.no_grad():
    generated_ids = model.generate(
        **inputs,
        max_new_tokens=max_new_tokens,
        do_sample=False,
    )

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(output)

示例输出:

unsafe
reasoning process
<answer>B. Real-World Crimes</answer>

第一行是二元判断,<answer> 包含来自默认分类体系或当前生效动态策略的最终风险类别。

注意事项

  • policy 会替换默认的风险规则。当启用动态策略时,请确保 <answer> 返回当前生效策略中的规则标题或 Safe。
  • 生产系统应处理格式异常的产出,例如无法解析的第一行、缺失的 <answer>,或不属于当前生效策略的类别。
  • 对于多模态输入,请确保图像路径对本地推理环境可访问。

风险类别

默认完整策略包含以下风险类别。当提供动态策略时,模型仅依据生效中的 policy 进行判断,而不会强制将每个案例归入默认类别。

A. 色情内容风险

  • 涉及露骨色情材料、性剥削或强制性行为的内容。

B. 现实世界犯罪与公共安全

  • 涉及暴力犯罪、武器、其他犯罪或公共安全威胁的内容。

C. 不道德行为

  • 涉及仇恨、骚扰、操纵、自残、令人不适的图像或有害虚假信息的内容。

D. 网络安全与信息操纵

  • 涉及数据泄露、黑客攻击、监控滥用、平台滥用或版权滥用的内容。

E. 智能体安全

  • 试图暴露系统提示词、内部策略或其他模型防护措施的内容。

F. 政治敏感内容

  • 涉及政治倡导、谣言、动乱、历史歪曲或攻击政治人物的内容。

G. 虐待动物

  • 涉及虐待动物或传播虐待动物行为的内容。

安全

  • 不符合任何现行风险类别的内容。

引用

@article{singguard2026,
  title={SingGuard: Policy-Adaptive Multimodal Safeguarding with Dynamic Reasoning},
  author={Ant Group},
  year={2026}
}

📄 许可证

本项目采用 Apache-2.0 许可证授权。

0.9B params

inclusionAI/SingGuard-0.8b 的模型树

包含 inclusionAI/SingGuard-0.8b 的合集

inclusionAI/SingGuard-0.8b 的论文

来源:蚂蚁 inclusionAI:HuggingFace 新模型 · huggingface.co