跳到正文
北京时间
原文
Hugging Face:Blog·· 2022-07-27精选AI 评分61

Hugging Face 用 TensorFlow 与 XLA 加速文本生成,部分场景提速超 100 倍

Faster Text Generation with TensorFlow and XLA

AI 导读

Hugging Face 宣布 🤗 transformers 的 TensorFlow 文本生成现可通过 XLA 编译加速,某些场景提速超过 100 倍,在多数基准中比 PyTorch 更快,最高快达 9 倍。

推荐理由

Hugging Face 官方讲解如何用 jit_compile=True 加速 TensorFlow 文本生成,给出编译缓存机制、padding 技巧和实测数据,方法可直接复用。

正文 · AI 翻译

TL;DR:在 🤗 transformers 上使用 TensorFlow 的文本生成现在可以通过 XLA 编译。它比以前快了多达 100 倍,并且甚至比 PyTorch 更快——请查看下面的 colab! Open In Colab

文本生成

随着大型语言模型质量的提升,我们对这些模型能力的期望也随之提高。特别是自 OpenAI 发布 GPT-2 以来,具备文本生成能力的模型一直备受关注。这有着充分的理由——这些模型可用于摘要、翻译,甚至在某些语言任务上展现出零样本学习能力。这篇博客文章将展示如何利用 TensorFlow 充分发挥这项技术的优势。

🤗 transformers 库最初是从 NLP 模型起步的,因此文本生成对我们来说至关重要也是理所当然的。这是 Hugging Face 民主化努力的一部分,旨在确保它易于获取、易于控制且高效。之前有一篇博客文章介绍了不同类型的文本生成。不过,下面还是对核心功能做一个快速回顾——如果你已经熟悉我们的 generate 函数并想直接跳到 TensorFlow 的具体细节,可以跳过这部分。

让我们从基础开始。文本生成可以是确定性的或随机性的,取决于 do_sample 标志。默认情况下它被设置为 False,使输出具有确定性,这也被称为贪婪解码。当它被设置为 True 时,也称为采样,输出将是随机性的,但你仍然可以通过 seed 参数获得可复现的结果(格式与无状态 TensorFlow 随机数生成相同)。一般来说,如果你想从模型中获得事实性信息,就需要确定性生成;如果你追求更具创造性的输出,就需要随机性生成。

# Requires transformers >= 4.21.0;
# Sampling outputs may differ, depending on your hardware.
from transformers import AutoTokenizer, TFAutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = TFAutoModelForCausalLM.from_pretrained("gpt2")
model.config.pad_token_id = model.config.eos_token_id
inputs = tokenizer(["TensorFlow is"], return_tensors="tf")

generated = model.generate(**inputs, do_sample=True, seed=(42, 0))
print("Sampling output: ", tokenizer.decode(generated[0]))
# > Sampling output: TensorFlow is a great learning platform for learning about
# data structure and structure in data science..

根据目标应用的不同,可能希望获得更长的输出。你可以通过 max_new_tokens 控制生成输出的长度,但请记住,更长的生成将需要更多资源。

generated = model.generate(
    **inputs, do_sample=True, seed=(42, 0), max_new_tokens=5
)
print("Limiting to 5 new tokens:", tokenizer.decode(generated[0]))
# > Limiting to 5 new tokens: TensorFlow is a great learning platform for
generated = model.generate(
    **inputs, do_sample=True, seed=(42, 0), max_new_tokens=30
)
print("Limiting to 30 new tokens:", tokenizer.decode(generated[0]))
# > Limiting to 30 new tokens: TensorFlow is a great learning platform for
# learning about data structure and structure in data science................

采样有几个可以调节以控制随机性的旋钮。最重要的是 temperature,它设定了输出的整体熵——低于 1.0 的值会优先采样概率更高的 token,而高于 1.0 的值则相反。将其设置为 0.0 会将行为退化为贪婪解码,而非常大的值则近似于均匀采样。

generated = model.generate(
    **inputs, do_sample=True, seed=(42, 0), temperature=0.7
)
print("Temperature 0.7: ", tokenizer.decode(generated[0]))
# > Temperature 0.7: TensorFlow is a great way to do things like this........
generated = model.generate(
    **inputs, do_sample=True, seed=(42, 0), temperature=1.5
)
print("Temperature 1.5: ", tokenizer.decode(generated[0]))
# > Temperature 1.5: TensorFlow is being developed for both Cython and Bamboo.
# On Bamboo...

与采样相反,贪婪解码在生成的每次迭代中总是选择最可能的 token。然而,这往往会导致次优的输出。你可以通过 num_beams 参数提高结果的质量。当它大于 1 时,会触发束搜索,它会持续探索高概率序列。这种探索以额外的资源和计算时间为代价。

generated = model.generate(**inputs, num_beams=2)
print("Beam Search output:", tokenizer.decode(generated[0]))
# > Beam Search output: TensorFlow is an open-source, open-source,
# distributed-source application framework for the

最后,在运行采样或束搜索时,你可以使用 num_return_sequences 返回多个序列。对于采样,这相当于从同一个输入提示多次运行生成;而对于束搜索,它会按降序返回得分最高的生成束。

generated = model.generate(**inputs, num_beams=2, num_return_sequences=2)
print(
    "All generated hypotheses:",
    "\n".join(tokenizer.decode(out) for out in generated)
)
# > All generated hypotheses: TensorFlow is an open-source, open-source,
# distributed-source application framework for the
# > TensorFlow is an open-source, open-source, distributed-source application
# framework that allows

如你所见,文本生成的基础操作很容易控制。不过,上面的示例并未涵盖所有选项,建议阅读 文档 了解高级用例。遗憾的是,当你在 TensorFlow 中运行 generate 时,可能会注意到执行需要一段时间。如果你的目标应用期望低延迟或大量输入提示,那么使用 TensorFlow 进行文本生成看起来是一项昂贵的努力。😬

别担心,本文的剩余部分旨在证明一行代码就能带来巨大改进。如果你想直接动手实践,这个 colab 提供了一个可以摆弄的交互式示例!

TensorFlow 与 XLA

XLA,即加速线性代数,最初是为加速 TensorFlow 模型而开发的编译器。如今,它也是 JAX 背后的编译器,甚至可以 与 PyTorch 一起使用。虽然“编译器”这个词对某些人来说可能令人生畏,但 XLA 在 TensorFlow 中使用起来很简单——它打包在 tensorflow 库中,并且可以通过任何创建图的函数中的 jit_compile 参数来触发。

对于熟悉 TensorFlow 1 的各位 🧓 来说,TensorFlow 图的概念很自然,因为那是当时唯一的操作模式。首先,你以声明式的方式定义操作来创建图。之后,你可以将输入通过图传递并观察输出。快速、高效,但调试起来很痛苦。随着 TensorFlow 2 的到来,引入了 Eager Execution 以及以命令式方式编写模型的能力——TensorFlow 团队在 他们的博客文章 中更详细地解释了其中的区别。

Hugging Face 在编写 TensorFlow 模型时考虑到了 Eager Execution。透明度是核心价值观,能够随时检查模型内部对此非常有益。然而,这确实意味着模型的某些使用无法开箱即用地受益于图模式的性能优势(例如,在调用 model(args) 时)。

幸运的是,TensorFlow 团队为我们这样的用户考虑到了 🥳!用 tf.function 包装包含 TensorFlow 代码的函数,会在你调用被包装的函数时尝试将其转换为图。如果你正在训练模型,调用 model.compile()(不带 run_eagerly=True)正是进行这种包装,这样你在调用 model.fit() 时就能受益于图模式。由于 tf.function 可以用于任何包含 TensorFlow 代码的函数,这意味着你可以将其用于模型推理之外的函数,从而创建一个单一的优化图。

现在你知道了如何创建 TensorFlow 图,用 XLA 编译它们就很简单了——只需将 jit_compile=True 作为参数添加到上述函数(tf.function 和 tf.keras.Model.compile)中。假设一切顺利(更多内容见下文),并且你使用的是 GPU 或 TPU,你会注意到第一次调用需要一段时间,但后续调用会快得多。下面是一个执行模型推理及其输出后处理的简单函数示例:

# Note: execution times are deeply dependent on hardware -- a 3090 was used here.
import tensorflow as tf
from transformers import AutoTokenizer, TFAutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = TFAutoModelForCausalLM.from_pretrained("gpt2")
inputs = tokenizer(["TensorFlow is"], return_tensors="tf")

def most_likely_next_token(inputs):
    model_output = model(inputs)
    return tf.argmax(model_output.logits[:, -1, :], axis=-1)

print("Calling regular function with TensorFlow code...")
most_likely_next_token(inputs)
# > Execution time -- 48.8 ms

只需一行,你就可以从上面的函数创建一个 XLA 加速的函数。

xla_most_likely_next_token = tf.function(most_likely_next_token, jit_compile=True)

print("Calling XLA function... (for the first time -- will be slow)")
xla_most_likely_next_token(inputs)
# > Execution time -- 3951.0 ms
print("Calling XLA function... (for the second time -- will be fast)")
xla_most_likely_next_token(inputs)
# > Execution time -- 1.6 ms

使用 TensorFlow 和 XLA 进行文本生成

与任何优化过程一样,天下没有免费的午餐——XLA 也不例外。从文本生成用户的角度来看,只有一个技术方面需要牢记。无需深入探究细节,以这种方式使用的 XLA 会在你调用时对 tf.function 进行即时(JIT)编译,这依赖于多态性。

当你以这种方式编译函数时,XLA 会跟踪每个张量的形状和类型,以及每个非张量函数输入的数据。该函数被编译为二进制文件,每次以相同的张量形状和类型(任意张量数据)以及相同的非张量参数调用时,编译后的函数都可以被重用。相反,如果你以不同的形状或类型的输入张量调用该函数,或者使用不同的非张量参数,那么将进行新的代价高昂的编译步骤。用一个简单的例子来总结:

# Note: execution times are deeply dependent on hardware -- a 3090 was used here.
import tensorflow as tf

@tf.function(jit_compile=True)
def max_plus_constant(tensor, scalar):
    return tf.math.reduce_max(tensor) + scalar

# Slow: XLA compilation will kick in, as it is the first call
max_plus_constant(tf.constant([0, 0, 0]), 1)
# > Execution time -- 520.4 ms

# Fast: Not the first call with this tensor shape, tensor type, and exact same
# non-tensor argument
max_plus_constant(tf.constant([1000, 0, -10]), 1)
# > Execution time -- 0.6 ms

# Slow: Different tensor type
max_plus_constant(tf.constant([0, 0, 0], dtype=tf.int64), 1)
# > Execution time -- 27.1 ms

# Slow: Different tensor shape
max_plus_constant(tf.constant([0, 0, 0, 0]), 1)
# > Execution time -- 25.5 ms

# Slow: Different non-tensor argument
max_plus_constant(tf.constant([0, 0, 0]), 2)
# > Execution time -- 24.9 ms

在实践中,对于文本生成,这仅仅意味着输入应填充到某个长度的倍数(因此可能的形状数量有限),并且使用不同的选项在第一次使用时会很慢。让我们看看当你天真地使用 XLA 调用生成时会发生什么。

# Note: execution times are deeply dependent on hardware -- a 3090 was used here.
import time
import tensorflow as tf
from transformers import AutoTokenizer, TFAutoModelForCausalLM

# Notice the new argument, `padding_side="left"` -- decoder-only models, which can
# be instantiated with TFAutoModelForCausalLM, should be left-padded, as they
# continue generating from the input prompt.
tokenizer = AutoTokenizer.from_pretrained(
    "gpt2", padding_side="left", pad_token="</s>"
)
model = TFAutoModelForCausalLM.from_pretrained("gpt2")
model.config.pad_token_id = model.config.eos_token_id
input_1 = ["TensorFlow is"]
input_2 = ["TensorFlow is a"]

# One line to create a XLA generation function
xla_generate = tf.function(model.generate, jit_compile=True)

# Calls XLA generation without padding
tokenized_input_1 = tokenizer(input_1, return_tensors="tf")  # length = 4
tokenized_input_2 = tokenizer(input_2, return_tensors="tf")  # length = 5
print(f"`tokenized_input_1` shape = {tokenized_input_1.input_ids.shape}")
print(f"`tokenized_input_2` shape = {tokenized_input_2.input_ids.shape}")

print("Calling XLA generation with tokenized_input_1...")
print("(will be slow as it is the first call)")
start = time.time_ns()
xla_generate(**tokenized_input_1)
end = time.time_ns()
print(f"Execution time -- {(end - start) / 1e6:.1f} ms\n")
# > Execution time -- 9565.1 ms

print("Calling XLA generation with tokenized_input_2...")
print("(has a different length = will trigger tracing again)")
start = time.time_ns()
xla_generate(**tokenized_input_2)
end = time.time_ns()
print(f"Execution time -- {(end - start) / 1e6:.1f} ms\n")
# > Execution time -- 6815.0 ms

哦不,那太慢了!如上所述,控制不同形状组合的一种解决方案是通过填充。分词器类有一个 pad_to_multiple_of 参数,可用于在接受任何输入长度和限制追踪之间取得平衡。

padding_kwargs = {"pad_to_multiple_of": 8, "padding": True}
tokenized_input_1_with_padding = tokenizer(
    input_1, return_tensors="tf", **padding_kwargs
)  # length = 8
tokenized_input_2_with_padding = tokenizer(
    input_2, return_tensors="tf", **padding_kwargs
)  # length = 8
print(
    "`tokenized_input_1_with_padding` shape = ",
    f"{tokenized_input_1_with_padding.input_ids.shape}"
)
print(
    "`tokenized_input_2_with_padding` shape = ",
    f"{tokenized_input_2_with_padding.input_ids.shape}"
)

print("Calling XLA generation with tokenized_input_1_with_padding...")
print("(slow, first time running with this length)")
start = time.time_ns()
xla_generate(**tokenized_input_1_with_padding)
end = time.time_ns()
print(f"Execution time -- {(end - start) / 1e6:.1f} ms\n")
# > Execution time -- 6815.4 ms

print("Calling XLA generation with tokenized_input_2_with_padding...")
print("(will be fast!)")
start = time.time_ns()
xla_generate(**tokenized_input_2_with_padding)
end = time.time_ns()
print(f"Execution time -- {(end - start) / 1e6:.1f} ms\n")
# > Execution time -- 19.3 ms

这样好多了,以这种方式执行的连续生成调用将比以前快几个数量级。请记住,在任何时候尝试新的生成选项都会触发追踪。

print("Calling XLA generation with the same input, but with new options...")
print("(slow again)")
start = time.time_ns()
xla_generate(**tokenized_input_1_with_padding, num_beams=2)
end = time.time_ns()
print(f"Execution time -- {(end - start) / 1e6:.1f} ms\n")
# > Execution time -- 9644.2 ms

从开发者的角度来看,依赖 XLA 意味着要意识到一些额外的细微差别。当数据结构的大小提前已知时,XLA 表现出色,例如在模型训练中。另一方面,当它们的维度无法确定或使用某些动态切片时,XLA 无法编译。现代文本生成实现是自回归的,其自然行为是扩展张量并在进行过程中突然中断某些操作——换句话说,默认情况下对 XLA 不友好。 我们重写了整个 TensorFlow 文本生成代码库,以向量化操作并使用带填充的固定大小结构。我们的 NLP 模型也被修改,以在存在填充结构的情况下正确使用其位置嵌入。除了 XLA 编译的可用性之外,结果对 TensorFlow 文本生成用户来说应该是不可见的。

基准测试与结论

上面你看到可以将 TensorFlow 函数转换为图并通过 XLA 编译加速它们。当前形式的文本生成只是一个自回归函数,在模型前向传递和一些后处理之间交替,每次迭代生成一个 token。通过 XLA 编译,整个过程得到优化,从而加快执行速度。但快多少呢?下面的 Gradio 演示包含一些基准测试,比较了 Hugging Face 的文本生成在多个 GPU 模型上针对两个主要 ML 框架 TensorFlow 和 PyTorch 的表现。

如果你探索这些结果,两个结论很快就会显现:

  1. 正如这篇博客文章一直在铺垫的,使用 XLA 时 TensorFlow 的文本生成速度要快得多。我们说的是在某些情况下超过 100 倍的加速,这真正展示了编译图的强大 🚀
  2. 在绝大多数情况下,使用 XLA 的 TensorFlow 文本生成是最快的选择,其中一些情况下速度提升高达 9 倍,打破了 PyTorch 是严肃 NLP 任务首选框架的迷思 💪

试试这个 colab,享受由 XLA 加速的文本生成的强大能力吧!

来源:Hugging Face:Blog · huggingface.co