LlamaIndex 工程师分享构建 LLM Agent 工具的实用技巧
Building Better Tools for LLM Agents
LlamaIndex 工程师在发布 LlamaHub Tools(含 Gmail、Google Calendar、Wolfram Alpha 等 15 个工具)后,总结编写 LLM Agent 工具的技巧。
作者结合 LlamaHub Tools 的开发实践,总结出输入容错、错误自恢复、日期函数等可直接套用的 Agent 工具编写方法。
在过去一个月里,我深入探索了大型语言模型(LLM)智能体的世界,并为智能体构建了 LlamaIndex 的工具库。上周,作为更广泛的 Data Agents 发布的一部分,我协助领导了 LlamaHub Tools 的工作。
立即探索我们的免费和付费计划。
在构建 LlamaHub Tools 的过程中,我收集了一些创建有效且易于使用的工具的技巧,并想分享我的一些想法。
LlamaHub Tools 的背景
LlamaHub Tools 允许像 ChatGPT 这样的 LLM 连接到 API,并代表用户执行创建、读取、更新和删除数据的操作。我们整理的工具示例包括起草和发送电子邮件、读取和创建 Google Calendar 邀请、搜索维基百科,而这只是我们发布时推出的 15 个工具中的一小部分。
工具抽象概述
那么 LlamaHub Tools 究竟是如何工作的呢?LlamaHub 工具抽象允许你轻松编写可被智能体理解和调用的 Python 函数。例如,与其试图让智能体进行复杂的数学运算,我们可以为智能体提供一个调用 Wolfram Alpha 并将结果返回给智能体的工具:
from llama_index.tools.base import BaseToolSpec
QUERY_URL_TMPL = "http://api.wolframalpha.com/v1/result?appid={app_id}&i={query}"
class WolframAlphaToolSpec(BaseToolSpec):
spec_functions = ["wolfram_alpha_query"]
def __init__(self, app_id: Optional[str] = None) -> None:
"""Initialize with parameters."""
self.token = app_id
def wolfram_alpha_query(self, query: str):
"""
Make a query to wolfram alpha about a mathematical or scientific problem.
Example inputs:
"(7 * 12 ^ 10) / 321"
"How many calories are there in a pound of strawberries"
Args:
query (str): The query to be passed to wolfram alpha.
"""
response = requests.get(QUERY_URL_TMPL.format(app_id=self.token, query=urllib.parse.quote_plus(query)))
return response.text上面的代码足以定义一个 LlamaIndex 工具,让智能体能够查询 Wolfram Alpha。再也不用在数学问题上猜测错误答案了!我们可以像这样初始化 Tool Spec 的实例:
# Initialize an instance of the Tool
wolfram_spec = WolframAlphaToolSpec(app_id="your-key")
# Convert the Tool Spec to a list of tools. In this case we just have one tool.
tools = wolfram_spec.to_tool_list()
# Convert the tool to an OpenAI function and inspect
print(tools[0].metadata.to_openai_function())以下是 print 语句清理后的输出:
{
'description': '
Make a query to wolfram alpha about a mathematical or scientific problem.
Example inputs:
"(7 * 12 ^ 10) / 321"
"How many calories are there in a pound of strawberries"
Args:
query (str): The query to be passed to wolfram alpha.',
'name': 'wolfram_alpha_query',
'parameters': {
'properties': {'query': {'title': 'Query', 'type': 'string'}},
'title': 'wolfram_alpha_query',
'type': 'object'
}
}我们可以看到,描述如何使用该工具的文档字符串被传递给了智能体。此外,参数、类型信息和函数名也被一并传递,让智能体清楚地知道如何使用这个函数。所有这些信息本质上充当了智能体理解该工具的提示。
继承自 BaseToolSpec 类意味着为智能体编写工具非常简单。事实上,上面的工具定义只有 9 行代码,不包括空白、导入和注释。我们可以轻松地让函数准备好供智能体使用,无需任何繁重的样板代码或修改。让我们看看如何将工具加载到 OpenAI 智能体中:
agent = OpenAIAgent.from_tools(tools, verbose=True)
agent.chat('What is (7 * 12 ^ 10) / 321')
""" OUTPUT:
=== Calling Function ===
Calling function: wolfram_alpha_query with args: {
"query": "(7 * 12 ^ 10) / 14"
}
Got output: 30958682112
========================
Response(response='The result of the expression (7 * 12 ^ 10) / 14 is 30,958,682,112.', source_nodes=[], metadata=None)
"""我们可以测试一下在不使用工具的情况下将此查询传递给 ChatGPT:
> 'What is (7 * 12 ^ 10) / 321'
"""
To calculate the expression (7 * 12^10) / 14, you need to follow the order of operations, which is parentheses, exponents, multiplication, and division (from left to right).
Step 1: Calculate the exponent 12^10.
12^10 = 619,173,642,24.
Step 2: Multiply 7 by the result from Step 1.
7 * 619,173,642,24 = 4,333,215,496,68.
Step 3: Divide the result from Step 2 by 14.
4,333,215,496,68 / 14 = 309,515,392,62.
Therefore, the result of the expression (7 * 12^10) / 14 is 309,515,392,62.
"""这个例子应该展示了为智能体编写新工具是多么容易。在本文的剩余部分,我将讨论我发现的一些编写更实用、更有效工具的技巧和诀窍。希望读完本文后,你会兴奋地想要编写并贡献一些自己的工具!
构建更好工具的技术
以下是一些编写更可用、更实用工具的策略,以尽量减少与智能体交互时的摩擦。并非所有策略都适用于每个工具,但通常下面至少有几个技巧会被证明是有价值的。
编写有用的工具提示
以下是一个工具的函数签名和文档字符串示例,智能体可以调用它来创建草稿邮件。
def create_draft(
self,
to: List[str],
subject: str,
message: str
) -> str:
"""Create and insert a draft email.
Print the returned draft's message and id.
Returns: Draft object, including draft id and message meta data.
Args:
to (List[str]): The email addresses to send the message to, eg ['adam@example.com']
subject (str): The subject for the event
message (str): The message for the event
"""这个提示利用了几种不同的模式来确保智能体能够有效地使用该工具:
- 给出函数及其用途的简洁描述
- 告知智能体该函数将返回什么数据
- 列出函数接受的参数,并附上描述和类型信息
- 为具有特定格式的参数提供示例值,例如 adam@example.com
工具提示应简洁,以免在上下文中占用过多长度,但也要足够详尽,使智能体能够正确使用工具而不出错。
让工具能够容忍部分输入
帮助智能体减少错误的一种方法是编写对其输入更加宽容的工具,例如在可以从其他地方推断出值时将输入设为可选。以起草电子邮件为例,但这次让我们考虑一个更新草稿电子邮件的工具:
def update_draft(
self,
draft_id: str,
to: Optional[List[str]] = None,
subject: Optional[str] = None,
message: Optional[str] = None,
) -> str:
"""Update a draft email.
Print the returned draft's message and id.
This function is required to be passed a draft_id that is obtained when creating messages
Returns: Draft object, including draft id and message meta data.
Args:
draft_id (str): the id of the draft to be updated
to (Optional[str]): The email addresses to send the message to
subject (Optional[str]): The subject for the event
message (Optional[str]): The message for the event
"""Gmail API 在更新草稿时要求提供上述所有值,然而仅使用 draft_id,我们就可以获取草稿的当前内容,并在智能体更新草稿时未提供这些值的情况下,将现有值用作默认值:
def update_draft(...):
...
draft = self.get_draft(draft_id)
headers = draft['message']['payload']['headers']
for header in headers:
if header['name'] == 'To' and not to:
to = header['value']
elif header['name'] == 'Subject' and not subject:
subject = header['value']
elif header['name'] == 'Message' and not message:
message = header['values']
...通过在 update_draft 函数中提供上述逻辑,智能体可以仅使用其中一个字段(以及 draft_id)来调用 update_draft,我们就可以按用户期望更新草稿。这意味着在更多情况下,智能体能够成功完成任务,而不是返回错误或需要询问更多信息。
验证输入与智能体错误处理
尽管在提示和宽容性方面做出了最大努力,我们仍可能遇到智能体以无法完成当前任务的方式调用工具的情况。然而,我们可以检测到这一点,并提示智能体自行恢复错误。
例如,在上面的 update_draft 示例中,如果智能体在没有 draft_id 的情况下调用该函数,我们该怎么办?我们可以简单地传递空值并从 Gmail API 库返回错误,但我们也可以检测到空的 draft_id 必然会导致错误,并改为向智能体返回一个提示:
def update_draft(...):
if draft_id == None:
return "You did not provide a draft id when calling this function. If you previously created or retrieved the draft, the id is available in context"现在,如果智能体在没有 draft_id 的情况下调用 update_draft,它就会意识到自己犯下的确切错误,并获得如何纠正该问题的指示。
根据我使用此工具的经验,智能体在收到此提示后通常会立即以正确的方式调用 update_draft 函数,或者如果没有可用的 draft_id,它会告知用户该问题并向用户询问 draft_id。无论哪种情况,都比崩溃或向用户返回库中不透明的错误要好得多。
提供与工具相关的简单函数
智能体可能会在那些对计算机来说原本很简单的函数上遇到困难。例如,在构建用于在 Google Calendar 中创建事件的工具时,用户可能会这样提示智能体:
在我的日历上创建一个事件,明天下午 4 点与 adam@example.com 讨论 Tools PR
你能看出问题吗?如果我们试着问 ChatGPT 今天是什么日子:
agent.chat('what day is it?')
# > I apologize for the confusion. As an AI language model, I don't have real-time data or access to the current date. My responses are based on the information I was last trained on, which is up until September 2021. To find out the current day, I recommend checking your device's clock, referring to a calendar, or checking an online source for the current date.智能体不知道当前日期,因此智能体要么错误地调用函数,为日期提供一个像 tomorrow 这样的字符串,要么根据其训练时间幻觉出一个过去的日期,要么将告知日期的负担推给用户。上述所有行为都会给用户带来摩擦和挫败感。
相反,在 Google Calendar 工具规范中,我们提供了一个简单的确定性函数,供智能体在需要获取日期时调用:
def get_date(self):
"""
A function to return todays date.
Call this before any other functions if you are unaware of the current date
"""
return datetime.date.today()现在,当智能体尝试处理上述提示时,它可以先调用该函数获取日期,然后按用户请求创建事件,推断出“明天”或“一周后”的日期。没有错误,没有猜测,也无需进一步与用户交互!
从执行变更的函数返回提示
有些函数会以某种方式对数据进行变更,以至于不清楚函数能向 agent 返回什么有用的数据。例如,在 Google Calendar 工具中,如果事件成功创建,把事件内容返回给 Agent 就没有意义,因为 agent 刚刚传入了所有信息,因此它已经在上下文中拥有了这些信息。
一般来说,对于专注于变更(创建、更新、删除)的函数,我们可以利用这些函数的返回值进一步提示 agent,从而帮助 Agent 更好地理解自己的操作。例如,从 Google Calendar create_event 工具中,我们可以这样做:
def create_event(...):
...
return 'Event created succesfully! You can move onto the next step.' 这有助于 agent 确认操作已成功,并鼓励它完成被提示要执行的操作,尤其是在创建 Google Calendar 事件只是多步指令中的一步时。我们仍然可以在这些提示中返回 id:
def create_event(...):
...
event = service.events().insert(...).execute()
return 'Event created with id {event.id}! You can move onto the next step.'将大型响应存储在索引中供 Agent 读取
在构建工具时已经提到过的一个考虑因素是 Agent 拥有的上下文窗口大小。目前,LLM 的上下文窗口往往在 4k-16k token 之间,但它当然也可能更大或更小。如果工具要返回的数据大小超过上下文窗口,Agent 将无法处理这些数据并报错。
在构建工具时已经提到过的一个考虑因素是 Agent 拥有的上下文窗口大小。目前,LLM 的上下文窗口往往在 4k-16k token 之间,但它当然也可能更大或更小。如果工具要返回的数据大小超过上下文窗口,Agent 将无法处理这些数据并报错。
在创建可能需要由 LoadAndSearchTool 包装的工具时,唯一需要考虑的是它们需要返回一个 LlamaIndex 文档列表。对于返回字符串的工具,要使其与 LoadAndSearchTool 兼容,你唯一需要做的修改是将其包装在一个文档和一个数组中:
from llama_index.readers.schema.base import Document
def large_text_response_function():
...
return result
def large_text_response_function():
...
return [Document(text=result)]验证 Agent 如何理解工具
在开发中调试工具的一个有用技巧是询问 Agent 关于它自己的工具:它有哪些可用工具、这些工具接受什么参数、这些参数代表什么,以及该工具用于什么目的。Agent 的响应有助于确定你的提示可能在哪些方面有所欠缺,或帮助查明为什么 Agent 无法成功使用你正在开发的工具。
一个调试 Google Calendar Tool Spec 的示例对话,假设你已经设置了类似于示例 notebook 的 Agent:
agent.chat('what tools do you have available')
# I have the following tools available
# 1. `load_data`: This tool allows me to load data from your calendar. It can retrieve a specified number of events starting from a given date.
# 2. `create_event`: This tool allows me to create an event on your calendar. You can provide details such as the title, description, location, start datetime, end datetime, and attendees.
# 3. `get_date`: This tool returns today's date. It can be used to ensure that the correct date is used when working with other tools.
# Please let me know if you would like more information about any of these tools or if you have any specific requests.
agent.chat('what arguments does the load_data tool take')
# 'The `load_data` tool takes the following arguments:
# 1. `number_of_results` (optional, default: 100): Specifies the number of events to return from your calendar.
# 2. `start_date` (optional): Specifies the start date from which to retrieve events. It can be provided as a string in ISO format (e.g., "2023-07-05") or as a `datetime.date` object.
# You can use these arguments to customize the number of events and the start date for retrieving data from your calendar.结语
为 Agent 构建工具要求你批判性地思考用户将如何与 Agent 交互,并尝试预见任何可能的使用模式。构建经过深思熟虑、能够对错误和问题具有韧性的工具,需要仔细考虑并测试可能的使用场景。我希望这些反思和技巧能帮助你为 Agent 创建新工具,别忘了在 LlamaHub 上分享你的工具。
来源:LlamaIndex:产品、工程与评测 · llamaindex.ai