跳到正文
北京时间
原文
SiliconFlow· @SiliconFlowAI · X·· 2026-05-12精选AI 评分74
AI 导读

信息的结构与呈现方式本身正成为AI智能层的关键。当前,让大语言模型以HTML格式输出,能提供比默认Markdown更丰富的视觉布局与交互性,是值得尝试的技巧。长远来看,人类虽偏好用音频输入,但视觉(图像/动画/视频)才是更理想的AI输出形式,因为大脑约三分之一皮层专司视觉处理。AI输出形态将沿“原始文本→Markdown→HTML→交互式神经视频/模拟”的路径演进,最终可能由扩散神经网络直接生成交互视频。同时,输入方式也需融合音频、文本、视频及手势等多模态交互。在人机输入输出深度融合方面,仍有巨大发展空间。

推荐理由

Karpathy 给的路线图从文本到 HTML 再到神经视频,其中第一步的‘让 LLM 输出 HTML’你今晚就能用上。未来交互形态的思考,值得产品经理细读。

正文 · AI 翻译

有时候,关键并不仅仅在于答案本身。

信息的组织与呈现方式,正逐渐成为智能层的一部分🧐

引用Andrej Karpathy@karpathy
顺便说一句,这真的非常有效:在你的查询末尾让你的 LLM“将你的回答结构化为 HTML”,然后在浏览器中查看生成的文件。我也曾成功让 LLM 把输出呈现为幻灯片等形式。 更一般地说,在我看来,音频是人类偏好的 AI 输入,而视觉(图像/动画/视频)则是他们偏好的输出。我们大脑中大约三分之一是一个专用于视觉的大规模并行处理器,它是信息进入大脑的十车道超级高速公路。随着 AI 的进步,我认为我们会看到一个利用这一点的演进过程: 1) 原始文本(阅读起来困难/费力) 2) markdown(粗体、斜体、标题、表格,稍微更护眼一些)<-- 当前默认 3) HTML(仍然带有底层代码的过程式形式,但在图形、布局甚至交互性上灵活得多)<-- 尚处早期,但正在形成新的良好默认 ...4,5,6,... n) 交互式神经视频/模拟 在我看来,这种外推(尽管相关技术目前还不存在)最终会走向某种由扩散神经网络直接生成的交互式视频。关于如何将精确/过程式的“软件 1.0”产物(例如交互式模拟)与神经产物(扩散网格)编织在一起,还有许多未解问题,但总体方向类似于最近爆火的 https://x.com/zan2434/status/2046982383430496444 在输入端也有必要的、尚待完成的改进。音频、文本或视频单独都不够,例如,我觉得需要指向和比划屏幕上的东西,就像你与一个实际坐在你身旁、和你共用电脑屏幕的人会做的所有事情一样。 TLDR 人类与 AI 之间的输入/输出心智融合正在进行中,还有很多工作要做,也有重大进展有待取得,远在我们一路跳到类似 neuralink 的脑机接口之类的东西之前。就当前阶段值得探索的东西而言,小贴士:试试要求 HTML。
原文

This works really well btw, at the end of your query ask your LLM to "structure your response as HTML", then view the generated file in your browser. I've also had some success asking the LLM to present its output as slideshows, etc. More generally, imo audio is the human-preferred input to AIs but vision (images/animations/video) is the preferred output from them. Around a ~third of our brains are a massively parallel processor dedicated to vision, it is the 10-lane superhighway of information into brain. As AI improves, I think we'll see a progression that takes advantage: 1) raw text (hard/effortful to read) 2) markdown (bold, italic, headings, tables, a bit easier on the eyes) <-- current default 3) HTML (still procedural with underlying code, but a lot more flexibility on the graphics, layout, even interactivity) <-- early but forming new good default ...4,5,6,... n) interactive neural videos/simulations Imo the extrapolation (though the technology doesn't exist just yet) ends in some kind of interactive videos generated directly by a diffusion neural net. Many open questions as to how exact/procedural "Software 1.0" artifacts (e.g. interactive simulations) may be woven together with neural artifacts (diffusion grids), but generally something in the direction of the recently viral https://x.com/zan2434/status/2046982383430496444 There are also improvements necessary and pending at the input. Audio nor text nor video alone are not enough, e.g. I feel a need to point/gesture to things on the screen, similar to all the things you would do with a person physically next to you and your computer screen. TLDR The input/output mind meld between humans and AIs is ongoing and there is a lot of work to do and significant progress to be made, way before jumping all the way into neuralink-esque BCIs and all that. For what's worth exploring at the current stage, hot tip try ask for HTML.

在 X 查看被引用的帖子

来源:SiliconFlow · x.com