AI 时代的“数据控制权”之争:Harness 成为新战场
The Harness Is the New Battleground
继 SaaS 之后,企业数据正通过 AI 使用轨迹(trajectories)流向模型厂商,引发数据主权担忧。纳德拉与帕兰提尔 CEO 同时警告数据泄露风险;7 月 13 日,安全研究员逆向 xAI 的 Grok Build 发现,零 AI 调用会话也会上传开发者代码库,xAI 已禁用该行为。未来 CIO 与 CEO 将要求零数据保留,厂商须承诺不访问企业数据。
作者把 Nadella 与 Karp 的同周表态和 Grok Build 上传代码事件串起来,指出 AI 时代企业数据可能经由 harness 变成厂商资产,数据留存条款或成采购核心。
The SaaS era broke precedent : for the first time, enterprises stored their data on vendors’ clouds, rather than servers in their buildings. Will this 20 year trend persist in the world of AI? The question is evocative & the debate of the moment.
Satya Nadella wrote over the weekend :
“You essentially pay for intelligence twice, once with money, & again with something even more valuable : the proprietary knowledge you must reveal to make that intelligence useful.”1
A week before, Alex Karp on CNBC’s Squawk Box said :
“[frontier labs] are stealing the weights & alpha of my business.”2
Karp & Nadella are both partners & competitors to the frontier labs. Their alignment on the same warning, in the same week, reflects a broader fear about data loss taking hold across enterprise software.
On July 13, a security researcher reverse-engineered xAI’s Grok Build binary & found that a session with zero AI calls had uploaded the developer’s codebase to xAI’s cloud.3 xAI has since disabled the behavior.
AI models need data to learn ; this isn’t new.
Google Analytics does this on the web. A user visits a candle shop’s landing page, clicks on the gift section, finds a coupon & checks out. The recommendation system improves for the next visitor.
Like the shopper at the candle shop, each time someone queries an AI, information is produced. This data is called a trajectory.
More data is always better for AI & AI labs pay for it. Startups produce trajectories by paying experts to use AI, capturing the data. Others train new AIs to synthetically create trajectories, mimicking users. Collectively, these companies generate roughly $10b in revenue & are some of the fastest growing startups ever.4
Unlike in SaaS, where this data lived in databases accessible only to the customer, trajectories can be fed back into a model to improve AI. The customer’s data can become part of a vendor’s intellectual property.
Could internal data commingle with trajectories? Or trade secrets? How do you answer customer support tickets? What is your brand identity? What are your employees paid?
All of it flows through a harness.
The harness is the software through which the user works with the AI, like Claude Cowork or Cursor.
The harnesses that win are the ones that maximize user productivity by being intelligent about how they marshal the AI.
CIOs & CEOs will demand zero data retention. Data fully deleted, not fully anonymized. The technologies around anonymization aren’t yet strong enough to guarantee it.
The trend needs to evolve & make the same guarantees software did : vendors have no access to enterprise data & don’t use it for their own purposes.
The last 20 years were the argument that vendors could be trusted with data. The next 20 will demand the same guarantees.
-
Satya Nadella, “The Reverse Information Paradox”, July 12, 2026. ↩︎
-
“Palantir’s Karp bashes token-based AI model as ‘completely wrong’”, CNBC Squawk Box, July 1, 2026. ↩︎
-
@hrkrshnn on X, July 13, 2026. ↩︎
-
@deedydas on X, “Every single startup selling AI Training Data (July 2026)”, July 12, 2026. ↩︎
来源:Tomer Tunguz 博客(VC 分析) · tomtunguz.com