跳到正文
北京时间
原文
elvis· @omarsar0 · X·· 2026-07-23精选AI 评分75
AI 导读

DAIR.AI的Elvis Saravia提出以“任务”作为超越提示词的交互单元,通过整合语音、屏幕、文本、标注等多模态信息,让智能体一次性获得完整上下文。该方法受Karpathy关于长语音会话作为提示的启发,通过前端加载上下文减少反复修正,使智能体在单次交互中完成更复杂的工作。

推荐理由

这篇文章把 Karpathy 的语音提示思路扩展成可落地的多模态任务方法,减少了代理交互的摩擦,做 agent 的可以试试,虽然简单但实用。

正文 · 原文

http://x.com/i/article/2079981292108582912

What Comes After the Prompt

Karpathy’s recent post about using long voice sessions as prompts helped me make sense of a prompting technique I now rely on often while building with agents. The visual that accompanies this post, From a prompt to a task, summarizes the idea in one picture.

For lack of a better term, I have been calling the unit a task. A task uses multimodal prompting to give an agent the instruction and as much relevant context as possible in one turn. It covers a larger unit of work than a single prompt, and it leaves behind a stored trace that can later become a reusable skill.

A task can include a long voice explanation, the current screen, precise text, annotations, transcriptions, images, and any other evidence that helps the agent understand the work. Each modality contributes something different. Voice carries reasoning, priorities, examples, and uncertainty. The screen gives the agent the current state and the environment where the work needs to happen. Annotations direct attention to specific details. Text preserves exact requirements, names, and constraints. Together, these signals give the agent a richer representation of the work.

The interaction

The experience feels closer to guiding an agent through a complex assignment than composing a conventional prompt. I front-load the context that would otherwise emerge across several turns, then give the agent room to complete more of the work in a single pass. In practice, I record a voice note while walking through the work, capture the relevant screen, mark it up with quick annotations, and paste in the exact text the agent needs.

The agent can still ask questions when important information is missing. In my experience, richer tasks reduce the repetitive back-and-forth where I restate context, point out the same details, or correct an assumption that could have been resolved from the beginning. A recent example was scheduling a post on a platform I rarely use. I recorded a short voice note with the goal and constraints, shared the screen with the scheduling page open, and annotated the fields that mattered. The agent completed the setup in one pass, and the usual follow-ups about which fields to fill and which copy to paste never happened.

This has also changed how I think about productivity with agents. A well-formed task gives me more confidence to hand off work and move to something else. That makes parallel work more practical because each agent needs less active supervision while it runs.

Why it works

Karpathy pointed out that LLMs are remarkably good at reconstructing intent from long, disorganized voice sessions. A ramble contains many weak signals about the goal, the constraints, the examples that matter, and the speaker’s uncertainty. The model can organize those signals into a cleaner representation of the request.

I am extending that idea with more modalities. The voice session provides the reasoning, while the screen, text, annotations, transcriptions, and images provide additional evidence. When one channel is noisy or incomplete, another channel can help resolve the ambiguity.

Complex agent tasks often fail at the boundaries between what I meant, what I explicitly said, and what the agent could observe. Multimodal prompting gives the model more opportunities to close those gaps before it begins the work.

Cost and payoff

This approach can look like overkill, and sometimes it is. A simple request still deserves a simple prompt. I use richer tasks when the work is long-running, when precision matters, when the agent needs to navigate an unfamiliar interface, or when a mistake would create several rounds of correction.

A multimodal task can also cost more because it contains more context. In my experience, that investment usually pays for itself. I can complete a larger unit of work per turn because the agent begins with more of the context it needs.

This is especially useful for browser use and computer use. The agent can see the environment, hear the reasoning behind the request, follow annotations that identify important elements, and use text for exact details. That combination helps the agent navigate unfamiliar interfaces.

Some of my current examples include scheduling posts on unfamiliar platforms, improving writing and editing, and refining the design of artifacts and web pages. These tasks involve many small decisions that are tedious to encode as a traditional prompt but easy to communicate while showing the work. In a design refinement task, the modalities map naturally. Voice explains what feels off about the layout and what the change should preserve. The screen shows the current state of the artifact. Annotations mark the specific spacing, components, or sections to adjust. Text supplies the exact copy and the constraints that should stay fixed.

From traces to skills

I store the traces from these tasks and review them for recurring patterns.

The useful patterns usually include the sequence of actions, the constraints I repeat, the quality checks I apply, and the corrections that consistently improve the result. Those patterns can be extracted into reusable skills so the next agent starts with a stronger workflow.

This connection to automation is important. A task gives me a practical unit that I can inspect, improve, and eventually place inside a larger loop. The richer initial trace helps me understand which parts can be automated reliably and where human guidance still adds value.

The process usually starts with a manual task. Repeated use produces traces, the traces reveal patterns, and the patterns become a reusable skill. Over time, the workflow requires less explanation because the important guidance has been captured.

If you are curious to learn more, I will be demoing, sharing, and writing more about this with our academy here: https://academy.dair.ai/

Toward omnimodels

Omnimodels, models built to consume voice, vision, images, and text natively, should make this style of interaction feel natural. We will be able to speak, show, point, type, and provide examples within the same session, while the model integrates those signals directly.

I feel like I am rehearsing for that interaction now. The current tools already make it possible, even if the experience still feels stitched together across voice, browser state, images, and text.

The term task is provisional, but the underlying idea has become clear through repeated use. Give the agent a richer trace of the work, let it reconstruct the intent, store what happened, and reuse the patterns that work.

This came from a practical need. I wanted fewer correction loops, stronger handoffs, and more dependable long-running agent workflows. Multimodal prompting has moved me steadily in that direction, and it has become my default way of handing agents real work.

来源:elvis · x.com