Jim Fan· @DrJimFan · X·· 2025-08-07精选
AI 导读
Qwen发布4B参数模型Qwen3-4B-Instruct-2507与Thinking-2507,支持256K上下文,分指令与推理双版本。作者指出这验证了"推理核心假设":推理仅需基础语言能力,无需千亿参数知识库,契合轻量级LLM OS理念——最小化模型体积,最大化依赖工具调用与知识检索。
推荐理由
Jim Fan 借 Qwen3-4B 提出「推理核心假设」,探讨极小模型作为 LLM OS 内核的边界
正文 · 原文
This may be a testament to the “Reasoning Core Hypothesis” - reasoning itself only needs a minimal level of linguistic competency, instead of giant knowledge bases in 100Bs of MoE parameters. It also plays well with Andrej’s LLM OS - a processor that’s as lightweight and fast as possible, and maximally relies on knowledge lookup, tool use, agentic flow, etc.
Now I’m curious - what’s the absolute smallest model we can squeeze that still functions as a competent LLM OS Kernel?
🚀 Introducing Qwen3-4B-Instruct-2507 & Qwen3-4B-Thinking-2507 — smarter, sharper, and 256K-ready! 🔹 Instruct: Boosted general skills, multilingual coverage, and long-context instruction following. 🔹 Thinking: Advanced reasoning in logic, math, science & code — built for expert-level tasks. Both models are more aligned, more capable, and more context-aware. Huggingface: https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507 https://huggingface.co/Qwen/Qwen3-4B-Thinking-2507 ModelScope: https://modelscope.cn/models/Qwen/Qwen3-4B-Instruct-2507 https://modelscope.cn/models/Qwen/Qwen3-4B-Thinking-2507在 X 查看被引用的帖子
来源:Jim Fan · x.com