跳到正文
北京时间
原文
LlamaIndex:产品、工程与评测·· 2023-11-15精选AI 评分64

LlamaIndex 0.9 发布:新增 IngestionPipeline 与多模态 RAG 模块

Announcing LlamaIndex 0.9

AI 导读

LlamaIndex 发布 0.9 大版本,引入 IngestionPipeline 数据摄取抽象,转换步骤自动缓存以加速重复运行。

推荐理由

官方梳理了 0.9 的数据摄取缓存、接口精简和 tokenizer 默认值变化,开发者可据此评估升级成本与迁移要点。

正文 · AI 翻译

我们勤奋的团队很高兴地宣布我们最新的重大版本,LlamaIndex 0.9!你现在就可以获取它:

立即探索我们的免费和付费计划。

pip install --upgrade llama_index

在 LlamaIndex v0.9 中,我们花时间优化了用户体验的几个关键方面,包括 token 计数、文本分割等!

作为其中的一部分,开发者应该注意一些新功能和对当前用法的细微更改:

  • 用于摄取和转换数据的新 IngestionPipline 概念
  • 数据摄取和转换现在会自动缓存
  • 更新了节点解析/文本分割/元数据提取模块的接口
  • 默认分词器的更改,以及自定义分词器
  • PyPi 的打包/安装更改(减少臃肿,新的安装选项)
  • 更可预测和一致的导入路径
  • 此外,测试版中:用于处理文本和图像的多模态 RAG 模块!

有疑问或顾虑?你可以在 GitHub 上报告问题或在我们的 Discord 上提问!

继续阅读以了解有关我们的新功能和更改的更多详细信息。

IngestionPipeline — 用于纯数据摄取的新抽象

有时,你只想从数据源中摄取和嵌入节点,例如,如果你的应用程序允许用户上传新数据。LlamaIndex V0.9 中的新概念是 IngestionPipepline。

一个 IngestionPipeline 使用了一个新概念 Transformations,应用于输入数据。

那么什么是 Transformation 呢?它可以是:

  • 文本分割器
  • 节点解析器
  • 元数据提取器
  • 嵌入模型

以下是基本使用模式的快速示例:

from llama_index import Document
from llama_index.embeddings import OpenAIEmbedding
from llama_index.text_splitter import SentenceSplitter
from llama_index.extractors import TitleExtractor
from llama_index.ingestion import IngestionPipeline, IngestionCache

pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=25, chunk_overlap=0),
        TitleExtractor(),
        OpenAIEmbedding(),
    ]
)
nodes = pipeline.run(documents=[Document.example()])

转换缓存

每次运行相同的 IngestionPipeline 对象时,它会缓存输入节点的哈希值 + 转换以及管道中每个转换的输出。

在后续运行中,如果存在缓存命中,则将跳过该转换并使用缓存的结果。这大大加快了重复运行的速度,并有助于在决定使用哪些转换时改善迭代时间。

以下是一个保存和加载本地缓存的示例:

from llama_index import Document
from llama_index.embeddings import OpenAIEmbedding
from llama_index.text_splitter import SentenceSplitter
from llama_index.extractors import TitleExtractor
from llama_index.ingestion import IngestionPipeline, IngestionCache

pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=25, chunk_overlap=0),
        TitleExtractor(),
        OpenAIEmbedding(),
    ]
)

nodes = pipeline.run(documents=[Document.example()])
nodes = pipeline.run(documents=[Document.example()])

pipeline.cache.persist("./test_cache.json")
new_cache = IngestionCache.from_persist_path("./test_cache.json")
new_pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=25, chunk_overlap=0),
        TitleExtractor(),
    ],
    cache=new_cache,
)

nodes = pipeline.run(documents=[Document.example()])

以下是另一个使用 Redis 作为缓存和 Qdrant 作为向量存储的示例。运行此操作将直接将节点插入你的向量存储,并在 Redis 中缓存每个转换步骤。

from llama_index import Document
from llama_index.embeddings import OpenAIEmbedding
from llama_index.text_splitter import SentenceSplitter
from llama_index.extractors import TitleExtractor
from llama_index.ingestion import IngestionPipeline, IngestionCache
from llama_index.ingestion.cache import RedisCache
from llama_index.vector_stores.qdrant import QdrantVectorStore

import qdrant_client
client = qdrant_client.QdrantClient(location=":memory:")
vector_store = QdrantVectorStore(client=client, collection_name="test_store")
pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=25, chunk_overlap=0),
        TitleExtractor(),
        OpenAIEmbedding(),
    ],
    cache=IngestionCache(cache=RedisCache(), collection="test_cache"),
    vector_store=vector_store,
)

pipeline.run(documents=[Document.example()])

from llama_index import VectorStoreIndex
index = VectorStoreIndex.from_vector_store(vector_store)

自定义转换

实现自定义转换很容易!让我们添加一个转换,在调用嵌入之前从文本中删除特殊字符。

转换的唯一实际要求是它们必须接受节点列表并返回节点列表。

import re
from llama_index import Document
from llama_index.embeddings import OpenAIEmbedding
from llama_index.text_splitter import SentenceSplitter
from llama_index.ingestion import IngestionPipeline
from llama_index.schema import TransformComponent

class TextCleaner(TransformComponent):
  def __call__(self, nodes, **kwargs):
    for node in nodes:
      node.text = re.sub(r'[^0-9A-Za-z ]', "", node.text)
    return nodes
pipeline = IngestionPipeline(
    transformations=[
        SentenceSplitter(chunk_size=25, chunk_overlap=0),
        TextCleaner(),
        OpenAIEmbedding(),
    ],
)
nodes = pipeline.run(documents=[Document.example()])

节点解析/文本分割 — 扁平化和简化的接口

我们使解析和分割文本的接口更加简洁。

之前:

from llama_index.node_parser import SimpleNodeParser
from llama_index.node_parser.extractors import (
	MetadataExtractor, TitleExtractor
) 
from llama_index.text_splitter import SentenceSplitter

node_parser = SimpleNodeParser(
  text_splitter=SentenceSplitter(chunk_size=512),
  metadata_extractor=MetadataExtractor(
  extractors=[TitleExtractor()]
 ),
)
nodes = node_parser.get_nodes_from_documents(documents)

之后:

from llama_index.text_splitter import SentenceSplitter
from llama_index.extractors import TitleExtractor 

node_parser = SentenceSplitter(chunk_size=512)
extractor = TitleExtractor()


nodes = node_parser(documents)
nodes = extractor(nodes)

以前,LlamaIndex 中的 NodeParser 对象变得极其臃肿,同时包含文本分割器和元数据提取器,这既给用户更改这些组件带来了痛苦,也给我们尝试维护和开发它们带来了痛苦。

在 V0.9 中,我们将整个接口扁平化为单个 TransformComponent 抽象,以便这些转换更易于设置、使用和自定义。

我们已尽力减少对用户的影响,但主要需要注意的是 SimpleNodeParser已被移除,其他节点解析器和文本分割器已提升为具有相同的功能,只是解析和分割技术不同。

任何旧的 SimpleNodeParser 导入都将重定向到最等效的模块 SentenceSplitter。

此外,包装对象 MetadataExtractor已被移除,以支持直接使用提取器。

所有相关完整文档见下方:

分词与Token计数——改进的默认设置与自定义

LlamaIndex 之前的一大痛点就是分词。许多组件使用不可配置的 gpt2 分词器进行Token计数,这给使用非 OpenAI 模型的用户带来了麻烦,甚至对 OpenAI 模型也需要一些类似这样的临时修补!

在 LlamaIndex V0.9 中,这个全局分词器现在可以配置,并默认使用 CL100K 分词器,以匹配我们默认的 GPT-3.5 LLM。

对分词器的唯一要求是它是一个可调用函数,接受一个字符串,并返回一个列表。

下面是一些配置示例:

from llama_index import set_global_tokenizer


import tiktoken
set_global_tokenizer(
  tiktoken.encoding_for_model("gpt-3.5-turbo").encode
)

from transformers import AutoTokenizer
set_global_tokenizer(
  AutoTokenizer.from_pretrained("HuggingFaceH4/zephyr-7b-beta").encode
)

此外,TokenCountingHandler 也获得了升级,具有更好的Token计数,并且在可用时直接使用来自 API 响应的Token计数。

打包——减少臃肿

为了现代化 LlamaIndex 的打包方式,V0.9 也带来了安装方面的变化。

这里最大的变化是 LangChain 现在是一个可选包,默认不会安装。

要将 LangChain 作为 llama-index 安装的一部分进行安装,你可以参考下面的示例。根据你的需求还有其他安装选项,我们也欢迎未来对 extras 的更多贡献。

# installs langchain
pip install llama-index[langchain]
 
# installs tools needed for running local models
pip install llama-index[local_models]

# installs tools needed for postgres
pip install llama-index[postgres]

# combinations!
pip isntall llama-index[local_models,postgres]

如果你之前在代码中导入了 langchain 模块,请相应地更新你的项目打包依赖。

导入路径——更一致、更可预测

我们正在对导入路径做两项更改:

  1. 我们从根级别移除了不常用的导入,以使导入 llama_index 更快
  2. 我们现在有了一致的策略,使“面向用户”的概念可以在一级模块中导入。
from llama_index.llms import OpenAI, ...
from llama_index.embeddings import OpenAIEmbedding, ...
from llama_index.prompts import PromptTemplate, ...
from llama_index.readers import SimpleDirectoryReader, ...
from llama_index.text_splitter import SentenceSplitter, ...
from llama_index.extractors import TitleExtractor, ...
from llama_index.vector_stores import SimpleVectorStore, ...

我们仍然在根级别暴露一些最常用的模块。

from llama_index import SimpleDirectoryReader, VectorStoreIndex, ...

多模态 RAG

鉴于最近 GPT-4V API 的发布,多模态用例比以往任何时候都更容易实现。

为了帮助用户使用这些功能,我们开始引入一些新模块来支持多模态 RAG 的用例:

  • 多模态 LLM(GPT-4V、Llava、Fuyu 等)
  • 用于联合图像-文本嵌入/检索的多模态嵌入(即 clip)
  • 多模态 RAG,结合索引和查询引擎

我们的文档中有一份多模态检索的完整指南。

感谢你们的所有支持!

作为一个开源项目,没有我们的数百名贡献者,我们就无法存在。我们非常感谢他们,以及全世界数十万 LlamaIndex 用户的支持。Discord 上见!

来源:LlamaIndex:产品、工程与评测 · llamaindex.ai