跳到正文
北京时间
原文
Hacker News:AI 热帖· damaru2·· 3 小时前精选AI 评分77

IMDEA Networks 论文:九款对话式 AI 服务向第三方泄露对话标题、提示词与截图

AI companies leak data to advertisers [pdf]

AI 导读

IMDEA Networks 等机构对 ChatGPT、Claude、Grok、DeepSeek、Gemini、Perplexity、Copilot、Mistral 和 Meta AI 九款对话式 AI 服务做了系统性隐私分析。

推荐理由

研究用静态与动态分析量化了主流对话式 AI 的第三方跟踪与对话内容泄露,读者可以借此了解自己的对话在同意与订阅不同设置下的暴露程度。

正文 · AI 翻译

译文尚不完整,完整内容请切换到原文。

像蝴蝶一样提示,像追踪器一样蜇人:对网页和移动端对话式 AI 代理的隐私分析

Guilherme Oliveira

IMDEA Networks

Miguel Sanchez

IMDEA Networks

Juan Manuel De Santa Olalla Gómez

IMDEA Networks

Roi S. Serna

IMDEA Networks/UC3M

Tautvydas Jackevicius

IMDEA Networks

Jorge Garcia-Herrero

Independent

Aniketh Girish

IMDEA Networks

Guillermo Suarez-Tangil

IMDEA Networks

Narseo Vallina-Rodriguez

IMDEA Networks

摘要

随着 OpenAI 等知名对话式 AI 提供商采用基于广告的商业模式,传统的网页和移动端追踪实践正扩展到对话式 AI 服务中 [15]。然而,尽管其采用率不断增长,对话式 AI 服务的追踪、数据共享和变现实践在很大程度上仍不透明,且受到研究人员、监管机构和公众相对有限的审视。在本文中,我们对九种知名对话式 AI 服务的网页和移动端部署进行了系统性隐私分析。结合静态和动态分析,我们研究了第三方广告与追踪服务(ATSes)的存在情况,刻画了其数据流,并评估了同意选择、订阅层级和访问控制机制如何影响对话向第三方的暴露。我们发现了对话式 AI 平台独有的隐私风险:多个提供商向第三方披露敏感的对话衍生制品——包括标题、提示和截图——且往往同时附带持久性用户标识符,从而实现用户归因。我们还发现,一些提供商在没有访问控制的情况下公开暴露对话永久链接,使追踪器能够读取整个对话。我们的发现揭示了传统追踪技术如何日益与 AI 介导的交互交织在一起,创造出可收集、推断和传播敏感用户及对话信息的新途径。为评估这些实践的更广泛影响,我们在 GDPR 和 ePrivacy 指令的背景下对其进行了分析。我们开展了一项负责任披露流程,涉及受影响的提供商和主管的欧洲数据保护机构。我们的结果表明,对话式 AI 服务引入了一种新型隐私攻击面,其中提供商生成的对话制品会受到追踪和公开暴露,凸显出需要更强有力的保障措施来规范 AI 介导的交互。

关键词

对话式 AI、LLM、隐私、移动端、网页、追踪器

本作品根据 Creative Commons Attribution 4.0 International License 授权。要查看该许可证的副本,请访问 https://creativecommons.org/licenses/by/4.0/ 或致信 Creative Commons, PO Box 1866, Mountain View, CA 94042, USA。 Proceedings on Privacy Enhancing Technologies YYYY(X), 1–18 © YYYY 版权归所有者/作者所有。https://doi.org/XXXXXXX.XXXXXXX

1 引言

大语言模型(LLM)的近期进展催生了 ChatGPT、Gemini 和 Claude 等对话式 AI 服务,这些服务能够支持持久交互、多模态处理和自主任务执行。随着这些服务在个人和专业活动中的采用不断增长 [43],服务提供商正在探索新的商业模式,以将其不断增长的用户群体变现。广告正成为其中一种模式,并可能将传统上与 Web 和移动平台相关的追踪与归因基础设施扩展到对话式 AI 中。例如,路透社报道称,OpenAI 与 Criteo 合作,于 2026 年初在美国为 ChatGPT 免费层用户开展了一项广告试点 [52]。然而,这些追踪技术的整合引发了独特的隐私担忧。与传统 Web 和移动应用不同,对话式 AI 服务会例行处理高度敏感的提示、上下文信息、行为模式、上传的文档以及持久交互历史,这些内容可能揭示用户生活和专业活动的私密方面。因此,将此类信息——尤其是在缺乏有意义的透明度或同意的情况下——披露给第三方追踪服务,可能使用户和组织面临重大的隐私风险。Jazlan 等人的先前工作考察了基于 Web 的对话式 AI 服务中第三方追踪的整合 [32],主要聚焦于识别追踪器并刻画其数据收集行为。然而,对话式 AI 服务独特的交互模式在其 Web 和移动客户端中引入了新的隐私风险:这些服务会生成对话衍生的产物——包括对话标识符、URL、标题、预览、提示、回复和交互元数据——这些内容可能被披露给第三方,或通过可公开访问的资源暴露。此外,这些暴露如何受到同意选择、隐私设置、订阅层级和访问控制机制的影响,在很大程度上仍未被探索。为弥补这一空白,我们研究了三个研究问题:

• RQ1:对话式 AI 服务在其 Web 和移动客户端中,在多大程度上整合了第三方追踪、分析、广告和归因基础设施?

• RQ2:对话式 AI 服务向第三方实体或通过可公开访问的资源暴露了哪些对话衍生产物和用户信息,以及这些披露会带来哪些隐私风险?

1Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al.

• RQ3:Cookie 同意选择、订阅层级、隐私设置和访问控制机制如何塑造对话式 AI 服务中信息的披露与可访问性?为回答这些问题,我们对九种主流对话式 AI 服务进行了系统性隐私分析,涵盖全部九家提供商的 Web 客户端以及其中八家提供 Android 移动应用的 Android 客户端。我们结合静态与动态分析,评估其在同意选择、订阅层级和访问控制配置方面的隐私实践。具体而言,我们做出以下贡献:(1)在所评估的服务中,我们识别出 44 个第三方组织,并观察到每个被评估的 AI 服务都至少集成了一个第三方广告或跟踪服务。我们进一步揭示了 Web 与 Android 客户端之间的显著差异,并识别出仅在用户明确接受非必要 Cookie 后才被激活的第三方服务,表明同意决定直接影响对话式 AI 平台的跟踪面(§5)。(2)我们发现了对话式 AI 服务特有的新型隐私风险。与主要观察浏览活动的传统跟踪系统不同,对话式 AI 平台生成的产物直接编码了用户交互。我们表明,6/9 的 Web 客户端和 3/8 的 Android 客户端将对话 URL、标题、提示词和截图披露给第三方服务,且往往伴随持久性用户标识符。我们进一步证明,这些披露的隐私影响在很大程度上受同意选择和分享功能的影响,若干提供商通过缺乏访问控制的公开可访问永久链接暴露完整对话。这些发现揭示了跟踪器和外部行为者可借以获取用户完整对话的新渠道(§6)。(3)我们根据欧盟数据保护法对观察到的实践进行了法律分析,评估跟踪器激活、同意机制、对话产物披露、身份关联实践以及公开可访问的对话资源与 GDPR 和 ePrivacy 指令的兼容性(§7)。我们的发现表明,将传统跟踪基础设施整合进对话式 AI 服务,为暴露敏感用户信息(包括对话衍生产物和公开可访问的对话)创造了新途径。更广泛地说,我们的结果挑战了将对话式 AI 服务视为用户与 AI 提供商之间机密交流的认知。相反,它们正日益融入更广泛的在线跟踪生态系统,为 AI 中介服务带来了重要的技术和监管挑战。

负责任披露。我们对所有已识别问题遵循了负责任披露流程,按照伦理考量部分所述通知了受影响的提供商和主管数据保护机构(DPA)。

2 背景

本节提供关于对话式 AI 服务及其与跟踪技术日益整合的背景(§2.1),以及 Web 和移动平台上常见部署的跟踪机制(§2.2)。

2.1 对话式智能体

对话式 AI 服务是基于 LLM 的系统,通过自然语言界面与用户交互,通常可通过网页和原生移动客户端访问。OpenAI 于 2022 年 11 月公开发布 ChatGPT,标志着 AI 行业的一个重大转折点,其用户数迅速达到数亿,并引发了一场在消费级和企业级生态系统中部署对话式 AI 平台的行业竞赛 [42]。此后,其他提供商也发布了竞争性服务,包括 Perplexity AI、Anthropic 的 Claude、Google 的 Gemini、Microsoft 的 Copilot、xAI 的 Grok 和 DeepSeek。尽管在架构和部署模式上存在差异,这些服务都具有若干共同特征。大多数对话式 AI 服务维护持久用户账户和交互历史,并集成搜索引擎、分析平台、遥测框架、广告基础设施和云托管 API 等外部服务。现代 AI 智能体还日益支持多模态能力,包括图像分析、语音交互、文档分析、浏览辅助和自主任务执行。这些服务的快速发展和普及放大了引入数据驱动变现模式的经济激励。近期行业动态和新闻稿表明,对话式 AI 服务正开始融入现有的广告和追踪生态系统,而非取代它。例如,Criteo 报告称,40% 的受访美国消费者已使用 AI 智能体进行产品发现和购物辅助 [15],而行业参与者越来越多地将智能体 AI 描述为定向广告和“智能体商务”的下一个机遇 [2]。同样,大型科技公司正在开发使 AI 智能体能够直接与商业平台和外部数字服务交互的基础设施,例如 Google 的 Universal Commerce Protocol(UCP),它促进 AI 驱动的商务和可互操作的智能体生态系统 [4]。

2.2 移动端和网页追踪

现代网页和移动服务与第三方广告和追踪服务深度交织,这些服务支持大规模的用户画像、个性化、归因和定向广告 [48]。在网页上,追踪器依赖 cookie、追踪像素、浏览器指纹和 cookie 同步等技术来收集浏览器元数据、交互事件、设备特征、网络信息和账户关联标识符,从而实现持久的跨站点识别和行为画像 [1, 17, 29, 34]。移动应用同样集成第三方 SDK [24, 48],这些 SDK 收集设备和行为数据,通常依赖平台支持的标识符,如 Android 广告 ID(AAID)和 Apple 的广告标识符(IDFA)[28, 53],以及哈希电子邮件地址(HEM)1 来支持跨设备追踪 [62, 64]。浏览器和移动操作系统都提供了限制此类追踪的机制。网页浏览器日益限制第三方 cookie、指纹识别和其他跨站点追踪

1哈希电子邮件地址(HEM)是从用户电子邮件地址派生的假名标识符,可在不传输明文电子邮件地址的情况下实现跨服务身份匹配 [25]。2

像蝴蝶一样提示,像追踪器一样蜇人:对网页和移动端对话式 AI 代理的隐私分析 隐私增强技术论文集 YYYY(X) 12345 CA CA

网页 Android 后端

用户 用户标识符

ATS 服务器

客户端 服务器端

对话式 AI 服务

提示 响应 代理 对话 标识符 聊天内容 嵌入的 第三方 SDK 对话产物 UserID、 DeviceID、 AnonID、 OrganizationID、 HEM ConvId、 ChatURL、 SharedID、 ShareURL 标题、提示、 截图

图 1:隐私风险。

技术,而内容拦截器和隐私增强扩展可以阻止对已知跟踪域的请求。移动平台则依赖应用沙箱、运行时权限以及对广告标识符和其他敏感资源访问的限制。然而,这些保护措施并不能消除第三方数据收集:嵌入式 SDK 在宿主应用的安全上下文中执行,并可能访问应用可用的数据和权限 [28]。跟踪机制和平台保护措施上的这些差异,促使我们分别分析对话式 AI 服务的网页和移动客户端,如 §4 所述。

3 对话式 AI 服务中的隐私风险

将传统广告和跟踪技术集成到对话式 AI 服务中会带来独特的隐私风险。与传统的网页和移动服务不同,这些服务通常会处理高度敏感和情境化的信息,包括自由形式的对话、行为交互、上传的文档、交互历史记录和持久用户画像。因此,嵌入式第三方跟踪技术不仅可以访问行为和设备信息,还可以访问编码了用户与 AI 提供商交互内容和情境的产物。我们考虑图 1 中所示的三个主要实体:用户、对话式 AI 服务(第一方)以及第三方 ATS,例如分析提供商、广告网络、崩溃报告服务和反欺诈产品。对话式 AI 服务包括客户端和服务器端组件。嵌入在网页或移动客户端中的第三方代码和库可能直接从客户端收集信息并传输到其云基础设施,而服务器端组件可能在用户设备边界之外向第三方披露信息。在这种情境下,当对话衍生产物被披露给第三方时,就会出现新的隐私风险。这些产物可能包括包含用户直接提供信息的提示、概括交互内容的对话标题、捕获对话情境的截图,或持久对话 URL(永久链接)。与传统的网页和移动服务上的跟踪方法不同,对话产物除了可能将模型响应暴露给潜在竞争对手外,还可以直接揭示用户与 AI 服务交互的内容、情境以及潜在的敏感性质。ChatGPT Claude Grok DeepSeek Gemini Perplexity Copilot 同意书和

订阅层级 层级(访客·免费·付费) 同意(忽略·拒绝·接受) 模式(默认·隐身) 第三方 分析 隐私 分析 §4.2 静态:Androguard 动态:插桩 AOSP §4.3 §6 §5 §4.1 §4.5 §4.4 第三方 分类 第一方 第三方 eTLD+1 3P API ATS Javascript Cookies HTTP/S 请求 9 项服务 已选 对话式 AI 服务选择 1 插桩与黑盒 分析方法 2 实验条件与输入 3 对话与 提示输入 敏感信息 我患有“X”病 AI 信息 .... 分析 4 与泄露 对话 产物 网站分析 Chrome 148 + DevTools (CDP) Android 应用分析 Pixel 3a / Android 12 数据流 MS Copilot Le Chat Meta AI

图 2:方法概述。

当此类信息与持久性的用户或设备标识符(如电子邮件哈希)一同被披露时,风险会被放大,使第三方能够将敏感的对话数据与单个用户关联起来,并可能跨会话或跨服务进行关联。此外,当底层资源缺乏充分的访问控制时,对话 URL 可能提供对额外内容的访问权限,从而可能将整个对话暴露给第三方。这些风险挑战了人们将对话式 AI 视为用户与 AI 提供商之间机密交互的认知。当与常规追踪基础设施相结合时,对话式 AI 生成的丰富且持久的产物开辟了新的渠道,敏感信息可借此被披露、关联到单个用户,或被第三方获取。

4 方法

图 2 概述了我们的研究方法,用以回答我们的三个研究问题。按照这一流程,我们首先描述代表性对话式 AI 服务的选择(§4.1),随后介绍针对其网页端(§4.2)和 Android 客户端(§4.3)的插桩与黑盒分析方法,我们如何评估同意选择和订阅层级的影响(§4.4),以及用于在对话式 AI 服务上触发系统性且可复现行为的对话输入(§4.5)。所有实验均于 2026 年 5 月在西班牙进行。

4.1 对话式 AI 服务选择

我们分析了一组同时支持网页端和 Android 客户端的知名对话式 AI 服务。我们并非追求广度最大化,而是对一小部分具有代表性、占据相当大市场份额的 AI 服务进行深入且系统的分析。由于缺乏各提供商可靠的市占率数据,我们使用客观的流行度代理指标来选择那些拥有庞大用户基础的服务,包括针对网页服务的 Tranco 排名 [35] 以及移动端的 Google Play 累计安装量。表 1 总结了本研究纳入的九项服务及其提供商、网页端和移动端实现,以及指导我们选择的流行度指标。所有这些

3 隐私增强技术会议论文集 YYYY(X) Oliveira 等

表 1:本研究分析的网页端和移动端对话式 AI 服务,按 Tranco 排名域名排序(2026 年 5 月)。

服务提供商 网站域名 Android 包名 Tranco 排名 Play 商店安装量 ChatGPT OpenAI chatgpt.com com.openai.chatgpt 48 >1B Claude Anthropic claude.ai com.anthropic.claude 617 >10M Grok xAI grok.com ai.x.grok 956 >100M DeepSeek DeepSeek chat.deepseek.com com.deepseek.chat 1,196 >50M Perplexity Perplexity AI perplexity.ai ai.perplexity.app.android 1,249 >100M Gemini Google gemini.google.com com.google.android.apps.bard 6,045 >1B MS Copilot Microsoft copilot.com com.microsoft.copilot 11,011 >50M Mistral (Le Chat) Mistral AI chat.mistral.ai ai.mistral.chat 12,039 >1M Meta AI Meta meta.ai com.facebook.stella 13,701 >50M

这些服务在 2026 年 5 月的 Tranco 2 中仍位列前 14K 服务。其中三个进入前 1K:ChatGPT(前 48)、Claude(前 617)和 Grok(前 956)。在移动端,它们的所有移动应用累计安装量均至少为 1M,其中两个(Gemini 和 ChatGPT)的累计安装量超过 1B。

4.2 网站分析

我们使用 Google Chrome(v148.0.7778.167)开发者工具(CDP)分析基于网络的对话式 AI 服务中跟踪器的存在及其数据收集实践,并保存生成的 HAR 跟踪记录以供后续分析。此设置能够观察用户交互期间生成的 HTTP(S) 请求、JavaScript 执行、浏览器存储访问、cookie、跟踪像素以及其他客户端跟踪机制。然后,我们检查与端点的通信,包括请求参数、请求体、cookie、浏览器存储条目和协议元数据,以识别以下内容的传输:(i) 用户标识符(例如账户 ID 和电子邮件地址);以及 (ii) 对话特定元数据、对话内容和标识符(例如聊天标识符和分享链接)。我们还搜索此类数据的转换表示形式,包括 Base64 编码以及用于生成哈希电子邮件地址的常见哈希算法(SHA-256、SHA-1 和 MD5)。我们检查浏览器存储机制并跨会话监控稳定 ID,以识别持久标识符和跟踪产物。最后,所有观察到的披露都映射到其对应的接收方域名,并与触发它们的实验配置相关联。所有会话均由一名研究人员手动进行,涵盖身份验证、引导流程和预定义的对话交流,详见 §4.5。实验在订阅层级、平台配置和隐私条件下重复进行,详见 §4.4,以评估同意场景(接受或拒绝非必要 cookie)和订阅层级(访客、免费和高级账户)的影响。试点实验显示,每种配置的跟踪行为都具有高度确定性,因此每种配置仅分析一次。

第三方域名分类。跟踪域名的标注由一位拥有十多年研究经验的合著者手动完成,采用保守标准以避免过度报告。我们使用公开可用的阻止列表和跟踪器情报来源来区分第一方和第三方域名,

2https://tranco-list.eu/list/3Q25L/1000000

包括 uBlock Origin [59] 和 whotracks.me [9],并借助 DNS 查询和证书透明度日志 [7] 的支持。由于一些对话式 AI 提供商同时运营广告、分析和云基础设施(例如 Google、Microsoft 和 Meta),仅考虑企业所有权可能会掩盖追踪关系。因此,我们将提供商自有的广告、分析和遥测端点归类为第三方广告与追踪服务(ATSes)。这些服务可能通过实时竞价请求促进与其他广告技术方的数据共享,甚至像 xAI 和 X Corp 那样以独立法律实体运营。这种方法为所有服务的追踪实践提供了一致的比较。

4.3 Android 应用分析

我们通过结合静态和动态技术来分析对话式 AI Android 应用,以最大化行为覆盖范围:

静态分析。我们使用 Androguard [16] 反编译每个应用的 APK。我们解析 AndroidManifest.xml 以枚举与追踪相关的已声明敏感权限(例如 AD_ID、READ_PHONE_STATE、ACCESS_FINE_LOCATION)。我们通过提取包命名空间,并按照先前工作实践 [24, 28, 64] 将应用的包名(例如 com.company.app)映射到其对应的 eTLD+1(company.com),来识别嵌入的第三方 SDK。所有权与应用 eTLD+1 不匹配的包和所联系域名被视为第三方组件 [48, 53]。然后,我们使用公开 SDK 文档和先前工作映射 [28, 45],以及 §4.2 中描述的第三方分类方法,手动匹配这些包和域名。

动态分析。我们在搭载 Android 12 构建的已插桩 Google Pixel 3a 上执行每个应用,该构建可透明地监控对受权限保护 API 的运行时访问、文件 I/O 操作以及所有出站网络流量;使用 mitmproxy [10] 和 Frida [46] 可实现同等覆盖。我们在系统层面观察对 TLS 套接字的读写,从而无需证书注入即可进行流量检查,且不会中断 TLS 握手,包括在证书固定应用中 [44, 47]。插桩追踪对敏感资源的访问,包括设备 ID(AAID、Android ID、IMEI、GSF ID、Boot ID)、硬件 ID(WiFi MAC 地址)、网络扫描数据(WiFi SSID、BSSID)以及账户关联 ID(例如电子邮件地址)。3 为补充静态

3每台设备都配置了假名 ID(电子邮件地址、电话号码),用于在每个平台上注册测试账户。由于每台设备的 ID 值已知,4像蝴蝶一样提示,像追踪器一样蜇人:Web 和移动对话式 AI 代理的隐私分析 隐私增强技术会议论文集 YYYY(X)

SDK 检测,我们对 Android Runtime 进行插桩,通过跟踪类链接器的 FindClass 方法来记录运行时加载的类,从而识别静态分析可能遗漏的混淆 SDK 组件。捕获的流量会自动解码常见编码(gzip、Base64)和 SDK 特定的混淆方法,并使用 §4.2 中描述的第三方端点分类方法进行解析,以提取字段名、值和目标端点。会话会被屏幕录制,以支持对数据流的事后验证。

4.4 同意表单和订阅层级

我们在不同的同意选择和订阅层级之间进行受控实验,从而评估这些因素如何塑造所观察到的隐私风险。我们的初步实验揭示了网页与移动端同意机制之间的差异。网页服务通常会展示 cookie 同意横幅,而大多数 Android 应用在启动时并不提供同等的同意机制。因此,我们评估每个客户端所支持的同意条件和隐私设置。对于网页客户端,我们分别进行显式接受或拒绝非必要 cookie 的运行,通过各服务提供的同意机制完成。这些交互是手动执行的,而非通过自动化手段,以准确捕获特定服务的同意流程和隐私控制。每次实验都从全新的浏览器配置文件开始,移除所有本地状态,以防止先前存储的标识符、同意选择、身份验证令牌或遥测痕迹影响后续测量。此外,在每个服务支持的情况下,我们还跨访客、免费和付费订阅层级进行受控交互。对于需要身份验证的层级,我们维护单独的账户,并在不同配置下执行等效的交互场景。Meta AI 和 DeepSeek 不提供高级层级,也都不提供访客层级。对于访客和付费配置,我们接受所有 cookie,以捕获流向第三方组织的数据流,并隔离与身份验证和订阅模式相关的差异。

4.5 对话与提示输入

为了捕获对话衍生痕迹连同用户标识符传播到第三方服务的动态证据,我们在每个平台上手动与每个服务交互,遵循预定义且可复现的交互协议。对于 §4.4 中描述的每种实验模式,我们进行一个聊天会话,包含若干旨在触发更广泛行为的提示。在服务支持的情况下,我们还会在单独的浏览器会话中生成并打开一个聊天分享链接,以评估过于宽松的访问控制是否会将共享对话暴露给第三方。虽然穷尽所有交互路径不可行,但这些提示在初步测量活动中经过迭代改进,以捕获具有代表性的行为。我们并不追求绝对完整性,而是使用丰富的健康相关内容来模拟真实用户,遵循基于角色的审计方法(§9)。具体而言,这些提示模拟一位用户咨询

我们可以自动在捕获的流量中搜索直接出现及其常见哈希变换(MD5、SHA-1、SHA-256)。

表 2:所测试的九种对话式 AI 服务中最常见的第三方组织。我们分别报告它们在网页和移动客户端中的存在情况,以及跨两种客户端类型(∩)和任一客户端类型(∪)的情况。条目按该组织在网页和移动客户端中同时存在的服务数量排序。在可能的情况下,组织按产品细分。图例:G# 仅网页,

H# 仅移动端,网页和移动端均有,若未联系则为空。

# 服务 按服务划分 组织/产品 Web 移动端 ∩∪ ChatGPT Claude Copilot DeepSeek Gemini Grok Meta AI Mistral Perplexity Google 8879G#H# Firebase 0808H#H#H#H#H#H#H#H# Search 8228G#G#G#G#G#G# Ads 7007G#G#G#G#G#G#G# Tag Manager 7007G#G#G#G#G#G#G# Accounts 5005G#G#G#G#G# Sentry 2424H#H# Datadog 3223G# Intercom 2103H#G#G# Meta 3113G#G#

该服务涉及医疗状况,引入敏感上下文,以在真实且隐私敏感的场景中评估平台行为。我们在§8.1中讨论了这种针对性方法的局限性;尽管如此,它为比较对话式AI服务之间的行为提供了一致的基线。

5 第三方服务分析

我们研究了第三方服务在对话式AI服务的Web和移动客户端中的结构性集成(§5.1),以及同意选择和订阅层级如何影响它们的存在(§5.2)。然后我们在§6中描述了传播给第三方组织的数据特征。

5.1 Web与移动端跟踪

每个对话式AI服务至少联系一个被归类为ATS的第三方组织。在我们的测量中,我们观察到124个不同的第三方域名,我们将其归因于44个组织,其中34个是ATS。图3展示了跨服务和客户端类型观察到的ATS生态系统。然而,第三方集成在Web和移动客户端之间存在显著差异。虽然11个ATS出现在两种客户端类型上,但15个仅出现在Web上,包括Google Tag Manager、TikTok以及OneTrust等同意管理平台。相比之下,8个ATS仅出现在移动客户端上,包括Braze。如表2所报告,Google产品在客户端类型中最为普遍,其次是Sentry、Meta、Datadog和Intercom。这些组织的存在反映了对话式AI服务广泛依赖第三方产品和服务,用于错误监控和可观测性、客户支持和参与,以及广告和分析等功能。Google

5Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al. Grok Claude Perplexity Copilot ChatGPT Mistral DeepSeek Gemini Meta AI Google Search Google Ads Google Tag Manager Google Accounts Datadog Meta Google User Content Intercom Sentry Apple Google Analytics Google APIs Doubleclick Dub Auth0 Inc. Appsflyer OneTrust ShuMei Sift Singular Sprig Stape Stripe TikTok Twitter Ads Twitter Analytics

(a) Web客户端。Perplexity ChatGPT Claude Copilot Mistral Grok DeepSeek Meta AI Google Firebase Sentry Google APIs RevenueCat Datadog Google Search Adjust Auth0 Inc. Appsflyer Fengkong Braze Meta Intercom OneTrust Sift Science Singular Stripe Twilio Segment Wingify (b) Android客户端。

图3:表示Web(左)和Android(右)客户端与ATS连接的二分图。图例:当用户接受非必要cookie或ToS(仅Android)时发生的交互;在拒绝cookie之前和之后均发生的交互(Web)。

拥有特别广泛的足迹,产品涵盖广告与分析(例如 Google Ads 和 Google Tag Manager)、搜索、身份验证(例如 accounts.google.com)、平台 API(例如 googleapis.com 下的子域名),以及 Firebase 等移动端专用的遥测服务。我们的分析还揭示了与 DeepSeek 相关的鲜为人知的中国追踪服务的存在,包括风控云(移动端,fp-it.fengkongcloud.com)和数美(Web 端,fp--it-acc.portal101.cn)。这些服务在学术文献中鲜有记载,也未出现在 WhoTracks.Me 等主流追踪器数据库中。然而,公开报告将它们与设备指纹识别和风险评分技术联系起来 [41]。

应用内浏览。移动客户端联系的所有不同端点中有 71% 源自 WebView(应用内浏览),而非原生字节码。4 对 WebView 的依赖程度在不同应用间差异很大,从 Grok 中 98% 的端点,到 Perplexity 中的 91%,再到 ChatGPT 中的 71%。相比之下,Copilot、Mistral 和 Meta 的大部分流量源自原生代码。虽然 WebView 便于将 Web 内容集成到原生应用中并简化开发,但它们也使移动应用内能够使用基于 Web 的追踪方法 [64]。例如,在 Grok 中,我们观察到 Google Ads、Google Tag Manager、TikTok Analytics、X/Twitter Analytics 和 Meta Pixel 的内容在 WebView 中加载。

5.2 同意与订阅层级的影响

同意决定会影响基于 Web 的客户端中追踪器的存在。相比之下,移动应用要求用户接受平台的服务条款(ToS)和隐私政策作为使用前提。因此,我们仅对 Web 客户端评估 §4.4 中描述的三种同意场景:(i) 忽略同意横幅(ignore),

4我们将每个出站流量归因于生成它的 Android UID,并将其分类为源自助手的原生包、Custom Tabs 或 WebView 进程。

(ii) 拒绝非必要 Cookie(reject all),以及 (iii) 接受所有 Cookie(accept all)。我们的结果表明,同意机制在不同提供商之间差异很大。一些服务允许用户在不做出明确选择的情况下继续交互(例如 Perplexity、Claude 和 Grok),而另一些则要求用户在访问平台前做出同意选择(例如 Gemini 和 Meta AI)。如图 4 所示,Mistral 要求用户在交互前完全接受其 ToS 和隐私政策。因此,所有连接都被标记为 accept all。附录 C 提供了观察到的同意表单示例。即使用户拒绝非必要 Cookie(即 reject all

场景下),我们仍观察到与第三方的连接。例如,Perplexity、DeepSeek、Gemini、Copilot、ChatGPT 和 Claude 在“全部拒绝”配置下会连接到 Google Ads,如图 3 中的虚线所示。接受所有非必要 cookie 会在 Claude、Perplexity 和 Grok 中激活额外的第三方。这些主要对应于知名的 ATS,包括 Meta、TikTok、Twitter Ads、DoubleClick 和 AppsFlyer。这些结果表明,一些提供商在用户同意后有条件地激活额外的广告和跟踪基础设施,如图 3 中的实线所示。忽略同意横幅不会导致连接到超出明确拒绝非必要 cookie 时所观察到的第三方域。订阅层级对联系的第三方集合影响相对有限。免费和高级账户在各项服务中表现出几乎相同的跟踪基础设施。一个例外是 Claude 的移动客户端,我们在免费层级中观察到 Intercom 和 Sentry,但在高级层级中没有。然而,此类差异可能源于在不同实验运行中变化的动态激活代码路径。

6像蝴蝶一样提示,像跟踪器一样蜇人:Web 和移动对话式 AI 代理的隐私分析 隐私增强技术会议论文集 YYYY(X)

表 3:跨 Web 和 Android 客户端的提供商与 ATS 之间的对话产物泄漏。由对话共享操作产生的共享 URL、提示和截图的传播在 §6.4 中研究。 提供商数量 ATS 数量 泄漏的产物 Web Android Web Android

对话 URL 5 0 9 0对话 ID 2 2 1 1用户提示 1 1 2 1对话标题 3 0 9 0对话截图 1 0 1 0共享对话 URL 5 0 9 0共享对话 ID 0 1 0 1

对话 ID 和共享对话 ID 计数仅涵盖标识符在没有完整 URL 的情况下泄漏的情况。

Web 上预置的跟踪面。观察到的网络流量仅捕获在特定动态测试期间激活的第三方子集。然而,对于基于 Web 的客户端,每个服务发出的内容安全策略(CSP)响应头(例如 script-src、connect-src、img-src 和 frame-src)列举了页面被授权联系或加载资源的其他外部源,从而揭示了可能在测试期间处于休眠或未触发的潜在关系,例如由于区域差异或执行上下文。CSP 头中包含的最普遍的第三方服务属于 Google Tag Manager(googletagmanager.com)、Google Analytics(google-analytics.com)和 Google Ads(googleadservices.com 和 doubleclick.net)。然而,如表 8 所示,CSP 策略通常包括广告行业中的其他知名参与者,例如 TikTok 和 Meta。

6 隐私分析

在 §5 中描述了第三方生态系统之后,我们现在研究敏感信息从 Web 和移动客户端向第三方的传播。

6.1 对话产物的传播

与传统网络追踪不同,对话式 AI 平台可能泄露编码了用户对话主题的痕迹:永久链接、提示词、对话标题和截图。表 3 总结了每种痕迹的第三方泄露数量。总体而言,6 个网页端和 3 个移动端对话式 AI 客户端在常规用户交互期间分别向 11 个和 2 个 ATS 泄露痕迹。这包括 Google Ads、TikTokinteractions 和 Twitter Analytics 等服务,详见表 4。

对话 URL。对话式 AI 服务会生成持久化的对话 URL,通常称为永久链接,用于唯一标识单个对话。对话永久链接是用于检索和管理对话历史的稳定 URL,直接绑定到特定对话。5 总体而言,我们观察到

5例如,在 Grok 的对话 URL 中:https://grok.com/c/0b3cc700-5a98-40f0-8e39-7311df76a700?rid=7b3e3eb2-0930-4490-b0ec-9f52fb3839c4,路径段

5 个网页端客户端向 9 个 ATS 披露对话 URL 或其关联的全局对话标识符。其中,60%(3/5)的网页端客户端默认执行此类披露,而其余客户端仅在用户明确接受非必要 cookie 后才这样做。ChatGPT 和 Claude 的网页端客户端还额外将全局唯一的对话 ID 作为独立参数泄露给 Datadog。尽管不如永久链接那样明确,但对话 ID 允许重建公开 URL。我们未在移动端客户端上观察到对话 URL 的披露。

对话标题。许多提供商会自动生成简短的对话标题,用于概括对话内容或目的,附录中的表 7 给出了示例。这些标题是 AI 生成的摘要,简洁地编码了关于底层对话主题、目的或意图的语义信息,从而揭示终端用户与 AI 系统讨论的敏感兴趣、意图、健康问题、财务状况、职业活动或其他个人话题。在受评估的服务中,33.3%(3/9)的网页端客户端向 9 个第三方泄露对话标题,包括 Meta、TikTok 和 Doubleclick。其中,88.9%(8/9)仅发生在用户接受非必要 cookie 时。在移动端版本上未观察到此类披露。

网页端 vs. 移动端。许多第三方同时出现在网页端和移动端助手中(§5),但它们收集的信息差异很大。网页端追踪器在页面的 JavaScript 上下文中执行,可以直接观察并记录浏览器和应用状态,包括聊天 URL、页面标题、对话标识符和追踪 cookie。相比之下,移动端客户端会收集对话 ID 等对话相关痕迹,但网页端版本中常见的面向用户的对话痕迹(如标题)则不存在。

同意与订阅层级的影响。拒绝非必要 cookie 可以减少向第三方传播信息,如表 4 所示。例如,在 Claude 中拒绝非必要 cookie 可阻止 Meta Pixel、Datadog 遥测以及向十一个广告平台的服务器端转发被激活。然而,在所有分析的免费层级服务中,即使拒绝了非必要 cookie,第三方跟踪器仍在 44.4%(4/9)的服务中收集数据。关于第三方的数据收集实践,我们未观察到免费与高级账户层级之间存在明显差异。

6.2 将对话关联到用户身份

当对话产物与用户或设备标识符一同传输时,其披露对隐私的侵入性显著增强,这些标识符如哈希电子邮件地址 [25] 或可重置的广告标识符(如 Android 广告 ID(AAID)[28, 64]),它们使第三方能够将这些交互与个人用户或长期跨平台行为画像关联起来 [62]。因此,我们分析对话式 AI 平台在多大程度上将对话产物与业界常用于广告归因、受众测量、分析和跨平台跟踪的标识符一同传输。

以及参数唯一标识该对话。然而,默认情况下,任何知道该 URL 的行为者均可公开读取该 URL,我们将在 §6.3 中进一步阐述。7

Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira 等

表 4:对话产物与用户标识符(例如电子邮件、哈希电子邮件和 cookie)同时传播给第三方。结果按平台、交互类型、同意状态和账户层级显示。

对话产物 用户身份 账户层级 LLM 服务平台 追踪器 永久链接 内容 用户 PII Cookie / SessionId 访客 免费 高级 ChatGPT ÜHDatadog ConvId 4–UserId 4, DeviceId 4SessionId 4ppp Claude ÜHDatadog ConvId , ShareId –AnonId 4, UserId 4––pp Perplexity ÜHWingify –Prompt ––p––ChatGPT ܀Datadog ConvId 4, ChatUrl –UserId 4, AnonId 4SessionId 4æææ Claude ܀Datadog ConvId 4–UserId 4, AnonId 4, OrganizationId 4SessionId 4–åå Claude ܀Intercom ConvId, ChatUrl –UserId 4, Email 4, OrganizationId 4, User Hash 4intercom-device-id 4–ææ Claude æ€Datadog ShareId, ShareUrl –AnonId SessionId å––Gemini ܀Google Analytics –Title –cid 4–ææ Gemini æ€Google Search ShareId, ShareUrl –––æ––Grok ܀DoubleClick ConvId, ChatUrl Title 4–em 4–åå Grok ܀Google Ads ConvId, ChatUrl Title 4–em 4–åå Grok ܀Google Search ConvId, ChatUrl Title 4–em 4–åå Grok ܀Google Tag Manager ConvId, ChatUrl Title 4–_fbp 4, _ttp 4, _dcid 4–åå Grok ܀Meta ConvId, ChatUrl Title –_fbp 4–åå Grok ܀TikTok ConvId, ChatUrl Title 4–_ttp 4–åå Grok ܀Twitter Analytics ConvId, ChatUrl Title 4–_twpid 4–åå Grok æ€Appsflyer ShareId, ShareUrl –––å––Grok æ€DoubleClick ShareId, ShareUrl Title ––å––Grok æ€Google Search ShareId, ShareUrl Title ––å––Grok æ€Google Tag Manager ShareId, ShareUrl Title –_dcid å––Grok æ€Meta ShareId, ShareUrl Title, Prompt –_fbp å––Grok æ€TikTok ShareId, ShareUrl Title, Prompt, Screenshot –_ttp å––Grok æ€Twitter Analytics ShareId, ShareUrl Title –_twpid å––Mistral ܀Intercom ConvId, ChatUrl Title UserId 4, Email 4, OrganizationId 4, User Hash 4intercom-device-id 4, intercom-id 4–åå Mistral æ€Intercom ShareId, ShareUrl Title –intercom-device-id, intercom-id å––Perplexity ܀Datadog ConvId 4, ChatUrl 4–UserId 4, AnonId 4, Email 4SessionId 4æææ Perplexity æ€Datadog ShareId, ShareUrl –AnonId –æ––

图例:€ 浏览器;H Android;Ü 正常对话期间披露;æ 访问共享对话时披露;4 在隐身模式下也会披露;å

仅在接受非必要 Cookie 后披露;æ 在接受和拒绝 Cookie 设置下均披露;p 无 Cookie 同意机制。Prompt 和 Screenshot 指对话中最近一次交互。在移动平台上,披露与 Cookie 同意设置无关。

表 5:收集对话产物及用户和设备 ID(包括追踪像素)的 ATS 数量。对于 Web 产品,我们还报告了仅在接受非必要 Cookie 后才出现的比例。

产品 平台 # ATSes % 全部接受

ChatGPT Web 1 0% Android 1 –Claude Web 2 50% Android 1 –Gemini Web 1 0% Grok Web 7 100% Mistral Web 1 100% Perplexity Web 1 0%

6.2.1 Web 追踪标识符。除了对话产物的披露之外,我们还考察了对话式 AI 服务通过追踪机制向第三方暴露 Web 标识符的情况,这些标识符可使第三方跨交互和浏览上下文识别或关联用户。总体而言,在观察到的向第三方传输 Web 标识符的行为中,77.3% 仅在用户接受非必要 Cookie 后发生。第三方追踪器同时收集多个标识符,使其能够将原本不同的用户身份和交互归并到同一档案下,甚至跨平台。我们观察到以下追踪技术:

• Web cookies and Tracking Pixels. Cookies and impression pixels remain the dominant web tracking mechanism. We ob-serve third-party cookies in 4 conversational AI services from 9 third parties, including Meta, TikTok, Twitter Analytics, and Intercom. Of those, 8 distinct cookies (e.g., _fbp , _ttp , _twpid )are transmitted alongside conversational artifacts. In the case of Grok Web, Meta Pixel alone collects the conversation ID, the chat title, the full conversation URL, the share ID, and the share URL, each linked to the user’s Meta identity by the synced _fbp

cookie. Meta documents the use of Pixel events for audience measurement and attribution, potentially aggregating them to users’ Meta accounts when one exists [38, 62].

• Email Hashes. Although often presented as privacy-preserving representations, hashed email addresses (HEMs) act as stable ID because third parties can independently compute identical hashes for known email addresses and use them for identity resolution, as pointed out by the USA Federal Trade Commission [25, 49]. Our results show that Perplexity transmits user’s HEMs to the marketing and analytics company Singular.

• Service Usernames and Account IDs. Four providers addition-ally transmit users’ usernames or account IDs on the AI services alongside conversational artifacts. Additionally, three providers share stable hash-based IDs to third parties like Intercom. Al-though we cannot determine the exact inputs used to generate these values, they remain consistent across sessions for the same account and therefore appear to function as persistent pseudony-mous IDs. For example, Mistral and Claude transmit such IDs to Intercom, while Grok sends a similar one to DoubleClick, Google Ads, and Google Search. These IDs may facilitate user identifica-tion and linkage across sessions even when direct account IDs are not disclosed.

• Cookie Syncing. We observe instances of cookie syncing [1] and server-side tracking [26] only when users accept non-essential cookies. Claude uses Segment Analytics proxied through a first-party domain (a-cdn.anthropic.com), loading a Conversion API configuration that forwards user events server-to-server from its infrastructure to eleven trackers (e.g., Facebook, LinkedIn, TikTok, Reddit, and Google Enhanced Conversions). This for-warding evades ad blockers and carries two shared IDs per event, which could enable data bridging across advertising platforms to

8Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents Proceedings on Privacy Enhancing Technologies YYYY(X)

a single session. Similarly, Grok embeds a Google Tag Manager container configured to route events through a server-side GTM instance (sGTM). A custom event transmits the conversation URL and chat topic title server-to-server to Meta Conversions API and TikTok Events API through the sGTM container, invisible to browsers and unblockable by ad blockers. The same payload carries both Meta’s _fbp and TikTok’s _ttp cookies, enabling ID bridging between Facebook and TikTok under a single identity.

• JS 指纹识别指标。虽然指纹识别超出了本工作的主要范围,但我们对实验中浏览器加载的 JavaScript 资源进行了探索性分析。遵循先前工作的发现 [1, 19, 30],我们在大多数提供商中搜索通常与浏览器指纹识别相关的 API。我们发现所有提供商提供的脚本中都调用了此类 API。尽管这些 API 的存在本身并不能确立主动的指纹识别行为,但它表明提供商具备潜在获取高熵设备特征的技术能力。这一观察与近期对 DeepSeek 的分析一致,该分析报告了通过 §5 中提到的风控云域名将指纹识别方法用于追踪和归因目的 [41]。

6.2.2 移动端实现。移动应用可以访问和暴露比浏览器环境通常可用的更广泛的标识符和敏感数据,从而实现更强的归因形式和跨平台用户识别。由于第三方 SDK 在宿主应用的进程中运行并继承其权限,用户授予应用的权限实际上可以被第三方搭便车利用 [24, 27, 28, 50]。在评估的对话式 AI 服务中,我们观察到以下向第三方服务的数据传播:

• 账户关联标识符和电子邮件地址。四个移动客户端向第三方披露账户关联标识符。两个传输电子邮件地址:Perplexity 传输给 RevenueCat,Grok 传输给 TikTok 和 Twitter Analytics。此外,Grok 将 HEM 传输给 TikTok 和 Twitter Analytics。

• 移动广告标识符(MAID)。三个移动客户端向第三方传输 Android 广告 ID(AAID)。我们发现 3 个不同的第三方接收 AAID:Adjust(Copilot)、AppsFlyer(Grok)和 Singular(Perplexity)。虽然 AAID 旨在支持广告归因和受众测量,但我们观察到它们与其他持久标识符一起传输的情况。特别是,Grok 将 AAID 与用户 ID 一起共享给 AppsFlyer,而 Perplexity 将 AAID 与用户和安装 ID 一起共享给 Singular [57]。这种组合使 ATS 能够将可重置的广告标识符与持久的账户或安装关联标识符相关联,从而破坏了可重置 ID 的隐私属性 [28, 40, 53]。即使用户重置了 AAID,安装级和账户级标识符仍然存在,并可用于在后续事件中重新绑定新分配的 AAID。

• SDK 特定 ID。Adjust、Auth0、Braze、Datadog、RevenueCat、Sentry、Sift Science 和 Singular 生成自己的专有 ID 用于用户识别、会话管理、归因和分析。在大多数情况下,这些 SDK 生成的 ID 独立于对话产物进行传输。然而,Claude 将对话 ID 转发给 Datadog,同时附带 Datadog 生成的

表 6:各提供商和账户层级的默认网站永久链接访问控制。

✘公开且无法退出 公开但可退出 G#仅限所有者 #不支持 额外控制:★白名单 4隐身模式 提供商 访客 免费 付费 控制 ChatGPT G#G#4G#4 Claude #G#4G#4 Grok #44 Deepseek #G## Perplexity ✘G#4G#4★ Gemini G#G#4G#4 Copilot G#G#4G#4 Mistral G#G#4G#4 Meta AI #G## 虽然所有受测服务都具备与所有人分享聊天并随后撤销此访问权限的功能,但部分提供商还提供了额外的访问控制机制,例如将特定用户加入白名单的能力(★)。当提供商提供隐身聊天(不保存在账户历史记录中)时,我们进一步用4标注相应层级。

以及与账户关联的 ID,从而在对话活动与多个专有分析 ID 之间建立起直接联系。

• 其他。Claude 将地理位置坐标和 Android ID(SSAID)[18] 都传输给 Sift Science,而 Meta AI 则将 Android ID 传输到其自己的服务器。虽然这些属性通常是为分析、安全或防欺诈目的而收集的,但当与其他 ID 结合时,它们可能增强用户识别与再识别的持久性和唯一性。我们还观察到移动端存在网页式跟踪 ID:Grok 基于 WebView 的集成会将 Meta 的 _fbp 和 TikTok 的 _ttp cookie 传输到与网页客户端相同的广告和分析端点。然而,与网页客户端不同,我们没有观察到对话 URL 或其他产物与这些 cookie 一同传输。总之,这些数据传播做法削弱了可重置 ID 本应在移动平台上提供的隐私保障。一次携带对话 URL 和 _fbp cookie 的 Meta Pixel 调用就能将聊天绑定到某个 Meta 个人资料;与 TikTok 或 Singular 共享的哈希电子邮件可与已知地址进行匹配 [25];而与持久安装 ID 一同传输的 AAID 则能在用户重置后依然存续。因此,对话活动会到达那些已经与长期和跨平台用户画像相关联的第三方。

6.3 对话永久链接的可访问性

当所引用的资源无需访问控制即可访问时,将持久对话 URL 或对话 ID 传播给第三方就变得尤为令人担忧。与传统的网页跟踪元数据不同,对话永久链接可以直接暴露用户交互内容,包括提示、生成的回复、兴趣、关切以及其他上下文信息,且如 §3 所讨论,这种暴露可能是持久性的。

9Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al.

为了评估这一风险,我们在未进行身份验证的情况下,于单独的隐私浏览会话中逐一打开每个对话 URL,从而系统地测试永久链接的可访问性。表 6 总结了在不同提供商和账户层级中观察到的访问控制机制。虽然一些提供商将对话限制为仅所有者可访问,除非明确共享——这一点将在 §6.4 中研究——但其他提供商则默认通过可公开访问的 URL 将其暴露,或依赖退出机制来限制访问。值得注意的是,Grok 采取了最宽松的立场:免费和高级层级的对话永久链接默认即可访问,并提供退出选项( )。Perplexity 始终将访客层级的对话设为公开(✘)。然而,在 2026 年 4 月 3 日,他们停止与 Meta 等第三方跟踪器共享 URL,这可能是对美国集体诉讼 [33] 的回应。在所有对话式 AI 提供商中,Web 和移动端实现之间的永久链接可访问性未观察到显著差异。我们并不试图确定这些设计决策背后的理由。然而,如果公开可访问性的目的是简化共享或提升可用性,我们的发现表明,这些选择通过扩大能够访问对话内容的实体集合,不必要地增加了用户面临的隐私风险。

6.4 对话共享

对话共享功能进一步放大了可访问性风险——所有服务都提供了生成可公开访问的对话 URL 的机制(见表 4 中的 æ)。我们观察到,当用户使用此功能时,所有对话式 AI 服务都会在无需身份验证的情况下暴露共享对话。在此,对话共享页面中存在的 9 个第三方可以完全查看这些对话。作为 Grok 对话共享功能的一部分,我们观察到 TikTok 收集对话截图,而 Meta 和 TikTok 都收集用户的最新提示。这些传输将对话的逐字部分暴露给第三方,如附录中的图 11 所示。

6.5 对话访问证据

为了调查传播的对话资源在上传后是否会被后续访问,我们部署了嵌入在对话中(通过在提示中直接共享 URL)和上传文件中(通过将 URL 嵌入上传的 .docx 和 PDF 文件)的 canary URL。我们还测试了一种变体,其中提示明确指示服务不要访问 URL。我们使用默认的 canarytokens.org [5],但也将其部署在自托管的 canary 基础设施 [58] 上,并使用通用域名,以降低 URL 被识别为 canary 的可能性。6

我们仅观察到 Grok 存在外部访问的证据,其中我们记录了 70 次不同的 canary 激活,时间跨度从初始提交后的数小时到数天不等。这些激活源自 70 个 IP 地址,跨越 14 个国家、4 大洲的 48 个 AS(自治系统)。虽然这些对话

6这一决定基于初步实验。一些 LLM 识别出与已知 canary token 域名相关的 URL 并拒绝访问。使用通用域名的自托管部署是为了避免这种情况。

在欧盟境内进行的连接中,65.7% 是由位于美国的机器触发的。与 Grok 在交互后反复激活不同,其余服务的访问通常仅限于最初的交互。对于 DeepSeek、Copilot、Mistral 和 Claude,一些金丝雀 URL 在提交时被激活过一次,请求来自 Amazon Web Services(AWS)和 Google Cloud Platform(GCP)等云服务提供商。Perplexity 反复访问金丝雀 URL——即使在被明确指示不要这样做时——来自一系列与其 Perplexity-User 爬虫相关联、托管在美国 AWS 基础设施上的 IP 地址。然而,未观察到访问不应被解释为不存在访问的证据,因为此类观察最终取决于我们的检测手段所提供的可见性以及下游系统访问此类资源所使用的具体机制。事实上,其他服务可能也支持对用户对话的服务器端访问,但正如 §8 所讨论的,我们无法观察到这种做法。然而,Grok 所观察到的反复检索表明,由对话衍生的资源在初始交互很久之后仍可能受到后续自动化访问的影响。

7 法律背景

本节讨论 §5 和 §6 中报告的所观察到的跟踪、数据共享和对话可访问性机制如何可能被置于 ePrivacy 指令和 GDPR 所确立的监管框架内,这两者管辖欧盟境内跟踪技术的使用和个人数据的处理。我们的目标不是提供确定性的法律评估或判定合规性,因为此类判定需要获取内部文档、合同安排、技术实现和组织措施,而这些并非公开可得。相反,我们识别出似乎与本研究中所观察到的做法相关的法律条款和监管指南,并讨论可能需要提供商、监管机构和研究人员进一步审查的潜在冲突领域。

7.1 ePrivacy 指令

According to Article 5(3) of the ePrivacy Directive (" ePD ") [22], the use of cookies for analytical, advertising, or profiling purposes —a category that unequivocally includes Meta and TikTok tracking pixels or Google Analytics cookies— requires prior informed user consent. The EDPB has clarified in detail how the prohibition under Article 5(3) ePD equally extends to more modern and sophisticated technologies than cookies [21], such as (i) URL and pixel tracking, (ii) collecting unique and persistent IDs, (iii) instructing the browser (through client-side code) to send such information to a server-side API. Additionally, compliance with applicable consent requirements includes: prohibition of pre-ticked boxes, prohibition of cookie walls except for an equivalent alternative, parity between accept and reject controls, and prohibition of dark patterns leading the user to an unreflective or manipulated acceptance.

7.2 GDPR Regulation (EU) 2016/679

When AI service providers disclose chat information to third parties, both (i) the disclosure and (ii) any subsequent access by such third

10 Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents Proceedings on Privacy Enhancing Technologies YYYY(X)

parties constitute processing activities requiring their own specific transparent information and legal grounds, provided that they are not strictly necessary for the performance of the contract between the provider and the user, and that they exceed the user’s legitimate expectations. Therefore, according to the GDPR [23], conversational AI services are legally obliged essentially: (1) To inform about their processing in a clear and under-standable manner. The user’s legitimate expectations would hardly include the possibility that their conversations could be accessible to third parties such as Meta or Google. Ope-nAI, Anthropic, Perplexity and xAI use generic and ambiguous expressions to refer to the user’s “user prompts” without specif-ically citing them when informing about third-party access: "user content " (OpenAI), “ conversations ” (Anthropic), “ service interaction info ” (Perplexity) and “ user info ” (xAI). Not only the EDPB, but also the CJEU in Meta Platforms Ireland (C-757/22, paragraph 54) [13] imposes literally on the data controller the obligation to inform the data subject linking for each partic-ular processing operation (i) the breakdown of personal data processed, (ii) its purpose, and (iii) the legal basis claimed. It is significant that even when OpenAI provides this information in such a structured and linked manner, it does not reference the practices observed in this study. (2) To have a legal basis for processing. The same applies to third parties accessing this data for their own purposes. This basis cannot be, as said, the “performance of the contract” between the user and the AI provider. It is highly debatable whether training the service with user interactions is “objectively neces-sary” for the provision of the service. We must bear in mind the constant assurances from dominant companies such as OpenAI or Anthropic regarding their policy of not training models on user interactions within the enterprise tier. What is not done with corporate users can hardly be “objectively necessary” for the same processing in the free tier. It seems even more difficult to legitimize as “necessary for the provision of the service” the access to third parties to the ti-tle and/or URL of the chats together with persistent IDs and tracking pixels for purposes unrelated to the operation of the tool, using adtech-specific mechanisms. The EDPB clarified this point in its binding decision 3/2022 against Meta IE (paragraph 118) [20] by indicating that if such processing was necessary, the contract terms should provide ( mutatis mutandis ) (i) the obligation of the provider to give third parties access to infor-mation about the chats, and (ii) contractual penalties in case of non-compliance with such “obligation”, elements that do not appear in the terms and conditions. This processing could only be carried out through explicit consent (as advocated by the EDPB) or legitimate interest. Even assuming the applicability of the legitimate interest, in the cases analyzed, the mandatory prior information on (i) processing and (ii) the possibility of exercising a right of opposition is conspicuously absent. Additionally, this data processing should be covered by some exception to the prohibition of processing special data cate-gories. It is a known and widespread practice among users (and encouraged by providers) that users make inquiries about their health, psychological state, intimate matters, etc, making the results of this study even more concerning. The longitudinal processing of these conversations over a significant period al-lows the inference of all kinds of special data categories even if not specifically protected (economic capacity, appetite for financial risk, vulnerability to compulsive purchasing).

Doctrine “SRB / Scania.” The “silver bullet” used by online ser-vices in this type of scenario is typically one of the following: (a)

anonymization or (b) pseudonymization under conditions that prevent the data recipient from re-identifying users without dispro-portionate efforts in terms of time, material resources, or personnel (SRB/Scania doctrine). The first thing to clarify is that these argu-ments would only serve (assuming hypothetically that they applied) to remedy the lack of a legal basis: the obligation to inform users about third-party access in the terms imposed by the CJEU in "EDPS vs SRB" (C-413/23 P, paragraph 110) [14] could have been breached. (1) Anonymization . The robustness and effectiveness of any ano-nymization process can only be assessed on a case-by-case basis, and it would be necessary to see the actual definition of anonymization and whether what has been done is sufficient. (2) Pseudonymization / SRB-Scania doctrine . It is highly un-likely that large platforms like Meta or Google could benefit from this doctrine, given their vast personal databases and their business model, which consists precisely and unequivocally of combining different data sources to better personalize targeted advertising. It should not be forgotten that the “insignificant risk of re-identification of the data subjects” must be assessed in light of the “context, purpose, and effects” CJEU Gesamtverband Scania (C-319/22, paragraph 70) [12] of the processing carried out by the party accessing the pseudonymized data.

8 Discussion

Our findings reveal that conversational AI services expose sub-stantially richer and sensitive information than conventional web and mobile services to tracking firms. Beyond collecting persistent IDs, several providers disseminate conversation-derived artifacts that encode the content, topic, or context of user interactions to third-party tracking services. When combined with stable IDs and permissive access-control mechanisms, these disclosures create new opportunities for third parties to associate sensitive conversational data with long-term user profiles and, in some cases, directly access the referenced conversation content. Therefore, our study raises broader questions about the exposure of conversational data, the effectiveness of existing privacy controls, and the accountability mechanisms governing these disclosures.

Conversation artifact leakage to the tracking ecosystem. Our results show that conversational AI has inherited the advertising and analytics infrastructure of the web and mobile platforms, bring-ing conversation data into existing tracking pipelines. Across 6 of 9 web services and 3 of 8 mobile apps, third parties receive conversa-tion artifacts like prompts, titles, and conversation permalinks (§6.1). These disclosures often occur alongside persistent IDs, including synced advertising cookies and HEMs (§6.2), making conversations linkable to long-term and cross-platform user profiles. Beyond the privacy implications, several third-party organizations, including Google and Meta, are embedded in competing AI products, raising

11 Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al.

questions about the collection of conversational data, including model responses, by direct competitors.

Privacy controls offer limited protection. The privacy controls available to users have limited impact on the underlying data flows. Even when non-essential cookies are rejected, 80.8% of third-party trackers remain active. Subscription tier is similarly ineffective, with free and premium accounts interacting with largely identical third parties. More fundamentally, the strongest controls are also least accessible to the users most exposed to tracking. On the web, only ChatGPT’s paid tier supports persistent consent preferences, yet roughly 95% of ChatGPT users remain on free tiers [51], and OpenAI has announced advertising for free and low-cost users [52]. On mobile, no equivalent to web cookie-consent mechanisms exist despite exposing similar tracking infrastructure as discussed in §5.

Scope and generalizability. While our evaluation focuses on leading consumer-facing services, the leakage vectors we observe stem from embedding standard web and mobile trackers directly into conversational interfaces. Consequently, these concerns extend beyond consumer B2C services to the broader ecosystem of LLM-powered web applications, custom customer-support chatbots, and third-party AI wrappers that rely on identical web and mobile analytics pipelines.

Exposure translates into access and liability. Beyond trans-mission, we also find evidence that the disclosed content is subse-quently retrieved. Our canary-token experiments indicate that at least some shared conversation artifacts are subsequently accessed. We observe repeated retrievals of conversations shared through Grok from infrastructure distributed across 14 countries, with 65.7% of accesses originating in the United States despite the conversa-tions being conducted in the EU (§6.5), adding data-sovereignty considerations to the underlying privacy concerns. However, the significance of the exposure does not depend on whether subse-quent access is directly observed. Conversational metadata reaches trackers through structured URL parameters and permissive sharing defaults (§6.3), reflecting deliberate design rather than incidental byproducts of page rendering. As the CJEU established in Fashion ID [11] (C-40/17, paragraph 78), enabling third-party access is itself a processing decision that triggers the provider’s obligations, re-gardless of whether the data is ultimately read (§7). The exposure, and the provider’s accountability for it, therefore arise from the design itself. Tooling and methodological constraints and research challenges, discussed next, limit our full visibility into this behav-ior and other server-side accesses. Addressing the challenges we faced is non-trivial, and we hope this work motivates the research community to tackle them.

8.1 Limitations

Our study covers nine conversational AI services, a subset of a broader, rapidly growing ecosystem. While selective in breadth, we focus on services with large active user bases. We exclude enterprise and governmental tiers, which providers market with distinct data-handling commitments whose validation we leave to future work. Our findings coverage is constrained to the limitations of black-box analysis techniques, including susceptibility to dynamic-analysis evasion, and we were unable to extract traces for Gemini’s mobile app. Furthermore, our evaluation focuses on discrete interaction sessions and does not capture cumulative privacy risks unique to long-lived conversations featuring persistent multi-turn memory. Nor does it cover unstudied scenarios such as native desktop clients, voice-first interfaces, or domain-specific vertical AI tools. The presence of a third party does not by itself imply that data is transmitted under all conditions only for advertising or tracking purposes, as some SDKs offer support for operational functions such as user attribution or in-app payments. Establishing processing pur-pose requires legal and contractual analysis beyond the scope of an empirical study. Large providers can also track users through their own first-party domains and server-side tracking, consistent with disclosures in their privacy policies, and our methodology can-not distinguish such tracking from functional traffic. We therefore report the presence of well-known first-party-tracking enablers such as Google Tag Manager. Finally, we measure from a fixed set of vantage points and user profiles, and practices in other geogra-phies, account states, or user contexts may differ. Moreover, as the ecosystem is rapidly evolving and subject to frequent updates, our findings represent a point-in-time lower bound on the extent of third-party integration and PII leakage in these services.

8.2 Mitigation

To mitigate some of these risks, both web browsers and mobile operating systems provide users with privacy and tracking controls. Privacy-enhancing browsers such as Brave block ads, analytics, and tracking resources at the network level by default [60]. How-ever, these could be defeated by advanced server-side and first-party tracking techniques. Similarly, mobile operating systems implement permission models that control app access to sensitive resources such as location, contacts, microphones, cameras, and MAIDs. Plat-forms such as iOS and Android also provide mechanisms to limit ad personalization and tracking through features such as App Tracking Transparency (ATT), advertising ID reset capabilities, and privacy dashboards that expose portions of application data access behavior. Building on these foundations, recent research explores auto-mated permission management to give users meaningful control over agent data access [68], an approach that consumer AI ser-vices could adopt to extend existing user controls over third-party trackers. A complementary direction may enforce website-provided policies through targeted sandboxing of browser-using agents [37], complementing user-facing permission controls with system-level confinement. Providers are also better positioned to act, for instance by not exposing conversation permalinks publicly by default and by withholding automatically generated titles and permalinks from third parties. Some providers have begun to expose dedicated con-trols. OpenAI, for example, recently introduced marketing-cookie controls [65], although the same update also permits using cookies to promote its products on other sites. From a user perspective, available mitigations include rejecting non-essential cookies and disabling chat sharing where platforms offer such controls. Never-theless, these measures remain partial and platform-dependent.

9 Related Work

A growing body of research focused on auditing the security of LLMs and their agentic extensions, uncovering vulnerabilities across

12 Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents Proceedings on Privacy Enhancing Technologies YYYY(X)

model training, information flows, tool integrations, and interac-tions with external services [8, 36, 56, 66, 69]. Our study focuses on privacy risks of generative AI in user-facing systems and fol-lows the dissemination of conversational data itself alongside user identifiers, including prompts, conversation titles, or chat IDs. Prior work relevant to this setting falls into three areas:

Third-party LLM Apps. An emerging app ecosystem layered atop foundation models exposes external services to the model via struc-tured tool interfaces. OpenAI’s GPTs and Anthropic’s MCP-based connectors are some of the few existing LLM app platforms. Prior work characterizes LLM apps as prompt-based wrappers combined with custom data-processing API integrations [31], examining their privacy posture without focusing on third-party integration [6]. Subsequent work audits OpenAI’s GPT app ecosystem, analyzing the natural-language specifications of GPT Actions to assess the data collection practices they declare [67]. Our work takes a more fundamental scope, studying the native AI layer of consumer ser-vices, which sits at the core of LLM app marketplaces.

Browser Agents and Extensions. Commercial agentic systems such as ChatGPT Atlas, Perplexity Comet, and AI-powered browser extensions like Sider and Monica act as external agents that drive the browser on behalf of their users. Prior studies enumerate the privacy risks of browser agents driving web sessions [55, 60], and audit popular GenAI browser extensions to measure data collection, processing, and sharing practices, as well as user profiling along demographic and interest attributes [61]. A complementary line of work characterizes browser and behavioral fingerprints of browsing agents [63], and evaluates same-origin policy enforcement across agentic browsers [54]. In contrast, we study assistants embedded directly in first-party services rather than installed as extensions, reaching a far broader user base by default. For instance, ChatGPT exceeds 1B installs just on the Play Store (cf. Table 1), compared to 5M for Monica, the most popular extension studied in [61].

First-Party Website and Mobile Services. A central contribution of our work is systematically studying and characterizing data leak-age from LLM services embedded in first-party websites and mobile applications, including conversational artifacts. Prior studies have examined third-party tracking in web-based AI services [32], pri-marily focusing on identifying trackers and characterizing their conventional data collection practices. Yet this work does not fully characterize the privacy risks arising from the generation, disclo-sure, and accessibility of conversation-derived information across both web and mobile AI services, covering their consent mecha-nisms and the impact of subscription tiers on tracker’s behavior. By jointly analyzing both web and mobile clients, conversation-derived artifacts, identity-linkage mechanisms, consent configurations, sub-scription tiers, and access-control policies, our work provides the first comprehensive assessment of privacy risks in mainstream conversational AI services. To probe these risks we adopt persona-based auditing with sensitive attributes, an established paradigm to uncover deep ad-targeting behavior [39].

10 Conclusions

This paper presented the first systematic measurement of third-party tracking in consumer conversational AI, analyzing nine promi-nent assistants across web and mobile clients under different con-sent states and subscription tiers. We showed that the tracking infrastructure of the web and mobile ecosystems—tracking pixels, SDKs, persistent identifiers, and server-side forwarding—has been carried into conversational AI, exposing conversation artifacts such as prompts, generated titles, and permalinks to third parties with advertising-based business models. These disclosures frequently occur alongside persistent identifiers, including advertising IDs, hashed email addresses, and tracking cookies, enabling conversa-tions to be linked to long-term user profiles. We further found that cookie consent and subscription tier provide limited protection, while permissive sharing defaults leave conversation permalinks publicly accessible. Using canary tokens, we confirmed that shared conversations are subsequently accessed from distributed third-party infrastructure. Together, these findings highlight tensions between current practices and obligations under the GDPR and ePrivacy Directive, exposing broader shortcomings in consent, ac-cess control, and transparency mechanisms and underscoring the need for stronger safeguards and accountability for conversational data flows.

Acknowledgments

This research was co-funded by RED2024-154240-T (EMACS) and by the European Union and the European Cybersecurity Compe-tence Centre under grant agreement No. 101309318 (REAL-PETS). G. Oliveira, JM. Santa Olalla Gómez, and T. Jackevicius acknowl-edge the financial support received through the CYBERACTION-ING project, co-funded by the European Union under the Digital Europe Programme (Grant Agreement No. 101123445), which sup-ported this research carried out as part of the Master’s Programme in International Cybersecurity and Cyberintelligence (MICAC). N. Vallina-Rodriguez and G. Suarez-Tangil have been appointed as 2019 Ramón y Cajal fellows (RYC2020-030316-I and RYC2020-029401-I, respectively) funded by MICIU/AEI/10.13039/501100011033 and the ESF Investing in your future. G. Suarez-Tangil and M. Sanchez’s work was supported by a 2025 Leonardo Grant for Scientific Re-search and Cultural Creation from the BBVA Foundation (SHIFT, reference LEO25-1-19720). The opinions, findings, and conclusions or recommendations expressed are those of the authors and do not necessarily reflect those of any of the funding agencies. The BBVA Foundation accepts no responsibility for the opinions, statements, and contents included in the project and/or the results thereof, which are entirely the responsibility of the authors.

Ethical Considerations

This research involves observing privacy risks in widely used con-versational AI services, so we took care to act responsibly through-out the process. We did not exploit any vulnerability, access other users’ conversations, or collect personal data from real users. All tests were carried out on accounts we owned only for research purposes, using a fixed prompt we authored ourselves, and traffic was captured locally through the browser’s developer tools.

13

Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al.

During our risk analysis, we determined that most of the issues reported in this study reflect intentional product-level practices rather than exploitable security vulnerabilities. Grok was the only service for which we identified a security concern that could be exploited by an external party, due to weak access controls on con-versation resources. This distinction guided our disclosure strategy. Our broader goal is to inform users, support regulatory oversight, and encourage providers to address the identified privacy risks and practices. Accordingly, we followed a responsible disclosure process for the exploitable security issue, providing xAI with an opportunity to investigate and remediate it prior to publication. We also notified relevant regulatory authorities of the broader privacy findings. The timeline of our disclosures was as follows:

• 23 March 2026 — We first noticed tracker activity while analyz-ing the network traffic of Perplexity and Grok.

• 6 April 2026 — We expanded the testing to cover major platforms in a systematic way, across different login states, cookie consent choices, account tiers, and privacy modes.

• 13 April 2026 — We notified our findings to the relevant Data Protection Authorities (DPAs) in the European Union and UK.

• 17 April 2026 — We notified xAI via their vulnerability disclosure email address about Grok’s conversation leaks due to their lack of access control mechanisms.

• 4 May 2026 — We released part of our findings public on the project website, LeakyLM.

• 27 May 2026 — The Spanish DPA, AEPD, cited our work request-ing to elevate the investigations for a plenary EDPB meeting to be held on June 6 2026 [3].

• 15 Aug 2026 OpenAI updated ChatGPT’s privacy policy, ex-plicitly mentioning third-party trackers; but we cannot confirm whether this action was directly influenced by of our findings.

• 10th Sept 2026 — Grok still uses publicly accessible permalinks. No official response has been received to date since our responsi-ble disclosure.

Open Science

To encourage reproducibility, we provide all research artifacts in https://github.com/guinucool/pbst2027.

AI Use

The authors used generative AI-based tools to revise the text, im-prove flow and clarity, correct typographical and grammatical er-rors, and create latex table structures. We have manually verified the integrity of this output. We do not use AI-based tools as part of our research methodology and data analysis pipeline, nor to generate figures. The literature review is of our own.

References

[1] Gunes Acar, Christian Eubank, Steven Englehardt, Marc Juarez, Arvind Narayanan, and Claudia Diaz. 2014. The web never forgets: Persistent tracking mechanisms in the wild. In Conference on Computer and Communications Security (CCS) .[2] AdExchanger. 2026. Daily News Roundup: OpenAI, Advertising, and AI Com-merce. https://www.adexchanger.com/daily-news-roundup/tuesday-03032026/ Accessed: 2026-05-31. [3] AEPD. 2026. La Agencia promueve ante las autoridades europeas de protección de datos que se estudie si algunos sistemas de IA permiten a terceros acceder a las conversaciones. https://www.aepd.es/prensa-y-comunicacion/notas-de-prensa/ la-agencia-promueve-ante-las-autoridades-europeas-proteccion-estudie-ia. Ac-cessed: 2026-09-7. [4] Google Developers Blog. 2025. Under the Hood: Universal Commerce Pro-tocol (UCP). https://developers.googleblog.com/under-the-hood-universal-commerce-protocol-ucp/ Accessed: 2026-05-31. [5] Canary 2026. Canarytokens. https://canarytokens.org/. Accessed: 2026-05-27. [6] Juan-Carlos Carrillo, Jose Luis Martin-Navarro, Rongjun Ma, and Jose Such. 2026. Personal Data Flows and Privacy Policy Traceability in Third-party LLM Apps in the GPT Ecosystem. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[7] Certificate Transparency. 2026. crt.sh: Certificate Transparency Search. https: //crt.sh/. Accessed: 2026-05-29. [8] Jeffrey Yang Fan Chiang, Seungjae Lee, Jia-Bin Huang, Furong Huang, and Yizheng Chen. 2025. Why are web ai agents more vulnerable than standalone llms? a security analysis. arXiv preprint arXiv:2502.20383 (2025). [9] Cliqz GmbH and Ghostery GmbH. 2026. WhoTracks.Me. https://whotracks.me Accessed: 2026-05-24. [10] Aldo Cortesi, Maximilian Hils, Thomas Kriechbaumer, and contributors. 2010–. mitmproxy: A free and open source interactive HTTPS proxy. https://mitmproxy. org/ [11] Court of Justice of the European Union. 2019. Fashion ID GmbH & Co. KG v Verbraucherzentrale NRW eV. https://eur-lex.europa.eu/legal-content/EN/TXT/ ?uri=CELEX:62017CJ0040 Case C-40/17, ECLI:EU:C:2019:629. [12] Court of Justice of the European Union. 2024. Gesamtverband Autoteile-Handel e.V. v Scania CV AB. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri= CELEX:62022CJ0319 Case C-319/22, ECLI:EU:C:2024:845. [13] Court of Justice of the European Union. 2024. Meta Platforms Ireland Ltd v Bundesverband der Verbraucherzentralen und Verbraucherverbände – Ver-braucherzentrale Bundesverband e.V. https://eur-lex.europa.eu/legal-content/ EN/TXT/?uri=CELEX:62022CJ0757 Case C-757/22, ECLI:EU:C:2024:598. [14] Court of Justice of the European Union. 2025. European Data Protec-tion Supervisor (EDPS) v Single Resolution Board (SRB). https://eur-lex. europa.eu/legal-content/EN/TXT/?uri=CELEX:62023CJ0413 Case C-413/23 P, ECLI:EU:C:2025:645. [15] Criteo. 2025. Agentic Commerce Is Emerging, Just Not the Way Most People Expect. https://www.criteo.com/blog/agentic-commerce-is-emerging-just-not-the-way-most-people-expect/ Accessed: 2026-05-31. [16] Anthony Desnos, Geoffroy Gueguen, and Sebastian Bachmann. 2015. Androguard: Reverse engineering, malware and goodware analysis of android applications... and more (ninja!). https://androguard.github.io/androguard/. Accessed: 2026-05-25. [17] Yana Dimova, Gunes Acar, Lukasz Olejnik, Wouter Joosen, and Tom Van Goethem. 2021. The cname of the game: Large-scale analysis of dns-based tracking evasion. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[18] Android Official Documentation. 2025. Settings.Secure. https://developer.android. com/reference/android/provider/Settings.Secure#ANDROID_ID Accessed: 2026-05-31. [19] Peter Eckersley. 2010. How Unique Is Your Web Browser?. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[20] European Data Protection Board. 2022. Binding Decision 3/2022 on the dis-pute submitted by the Irish Supervisory Authority on Meta Platforms Ireland Limited and its Facebook service (Art. 65 GDPR) . Technical Report. European Data Protection Board. https://www.edpb.europa.eu/system/files/2023-01/edpb_ bindingdecision_202203_ie_sa_meta_facebookservice_redacted_en.pdf Adopted pursuant to Article 65 GDPR. [21] European Data Protection Board. 2024. Guidelines 2/2023 on Technical Scope of Art. 5(3) of ePrivacy Directive . Technical Report. European Data Protec-tion Board. https://www.edpb.europa.eu/system/files/2024-10/edpb_guidelines_ 202302_technical_scope_art_53_eprivacydirective_v2_en_0.pdf [22] European Parliament and Council of the European Union. 2002. Directive 2002/58/EC concerning the processing of personal data and the protection of privacy in the electronic communications sector (ePrivacy Directive). Official Journal of the European Union, L 201, 31 July 2002, pp. 37–47. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02002L0058-20091219 Article 5(3), as amended by Directive 2009/136/EC. [23] European Parliament and Council of the European Union. 2016. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation). Official Journal of the European Union, L 119, pp. 1–88. https://eur-lex.europa.eu/eli/reg/2016/679/oj [24] Álvaro Feal, Julien Gamba, Juan Tapiador, Primal Wijesekera, Joel Reardon, Serge Egelman, and Narseo Vallina-Rodriguez. 2021. Don’t accept candy from strangers: An analysis of third-party mobile sdks. Data Protection and Privacy: Data Protection and Artificial Intelligence 13 (2021), 1. [25] Federal Trade Commission. 2024. No, hashing still doesn’t make your data anonymous. https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/07/ no-hashing-still-doesnt-make-your-data-anonymous. Accessed: 2026-05-31. 14 Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents Proceedings on Privacy Enhancing Technologies YYYY(X)

[26] Imane Fouad, Cristiana Santos, and Pierre Laperdrix. 2024. The Devil is in the Details: Detection, Measurement and Lawfulness of Server-Side Tracking on the Web. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[27] Julien Gamba, Álvaro Feal, Eduardo Blazquez, Vinuri Bandara, Abbas Razagh-panah, Juan Tapiador, and Narseo Vallina-Rodriguez. 2023. Mules and permission laundering in android: Dissecting custom permissions in the wild. IEEE Transac-tions on Dependable and Secure Computing 21, 4 (2023), 1801–1816. [28] Aniketh Girish, Joel Reardon, Juan Tapiador, Srdjan Matic, and Narseo Vallina-Rodriguez. 2025. Your Signal, Their Data: An Empirical Privacy Analysis of Wireless-scanning SDKs in Android. In Proceedings on Privacy Enhancing Tech-nologies (PoPETs) .[29] Alejandro Gómez-Boix, Pierre Laperdrix, and Benoit Baudry. 2018. Hiding in the crowd: an analysis of the effectiveness of browser fingerprinting at large scale. In Proceedings of the ACM Web Conference (WWW) .[30] Umar Iqbal, Steven Englehardt, and Zubair Shafiq. 2021. Fingerprinting the Fingerprinters: Learning to Detect Browser Fingerprinting Behaviors. In IEEE Symposium on Security and Privacy (S&P) .[31] Umar Iqbal, Tadayoshi Kohno, and Franziska Roesner. 2024. LLM platform security: Applying a systematic evaluation framework to OpenAI’s ChatGPT plugins. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society .[32] Muhammad Jazlan, Ethan Wang, Yash Vekaria, and Zubair Shafiq. 2026. Tracking Conversations: Measuring Content and Identity Exposure on AI Chatbots. arXiv preprint arXiv:2604.27438 (2026). [33] John Doe. 2026. Complaint for Damages and Demand for Jury Trial: Doe v. Perplexity AI, Inc. Complaint filed in the Superior Court of Califor-nia. https://cdn.arstechnica.net/wp-content/uploads/2026/04/Doe-v-Perplexity-Complaint-3-31-26.pdf Filed March 31, 2026. Available online. [34] Pierre Laperdrix, Gildas Avoine, Benoit Baudry, and Nick Nikiforakis. 2019. Morel-lian analysis for browsers: Making web authentication stronger with canvas fingerprinting. In Conference on Detection of Intrusions and Malware, and Vulner-ability Assessment (DIMVA) .[35] Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Ko-rczynski, and Wouter Joosen. 2019. Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation. In Network and Distributed System Security Symposium (NDSS) .[36] Shen Li, Liuyi Yao, Lan Zhang, and Yaliang Li. 2025. Safety Layers in Aligned Large Language Models: The Key to LLM Security. In International Conference on Learning Representations (ICLR) .[37] Luoxi Meng, Henry Feng, Ilia Shumailov, and Earlence Fernandes. 2025. cellmate: Sandboxing browser ai agents. arXiv preprint arXiv:2512.12594 (2025). [38] Meta for Developers. 2025. Meta Pixel Reference. https://web.archive.org/web/ 20250531104925/https://developers.facebook.com/docs/meta-pixel/reference/. [39] Maaz Bin Musa and Rishab Nithyanand. 2022. ATOM: Ad-network Tomography. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[40] Trung Tin Nguyen, Michael Backes, and Ben Stock. 2022. Freely given consent? studying consent notice of third-party tracking and its violations of gdpr in android apps. In Conference on Computer and Communications Security (CCS) .[41] NowSecure. 2025. NowSecure Uncovers Multiple Security and Privacy Flaws in DeepSeek iOS Mobile App. https://www.nowsecure.com/blog/2025/02/06/ nowsecure-uncovers-multiple-security-and-privacy-flaws-in-deepseek-ios-mobile-app/. Accessed: 2026-05-28. [42] OpenAI. 2022. Introducing ChatGPT. https://openai.com/blog/chatgpt Accessed: 2026-05-31. [43] OpenAI. 2025. How people are using ChatGPT. https://openai.com/index/how-people-are-using-chatgpt. Accessed: 2026-09-8. [44] Amogh Pradeep, Muhammad Talha Paracha, Protick Bhowmick, Ali Dava-nian, Abbas Razaghpanah, Taejoong Chung, Martina Lindorfer, Narseo Vallina-Rodriguez, Dave Levin, and David Choffnes. 2022. A comparative analysis of certificate pinning in Android & iOS. In Proceedings of the Internet Measurement Conference (IMC) .[45] Exodus Privacy. 2024. Homepage. https://exodus-privacy.eu.org/en/. Accessed: 2026-05-31. [46] Ole André Vadla Ravnås and contributors. 2014. Frida: Dynamic instrumentation toolkit for developers, reverse-engineers, and security researchers. https://frida. re/ [47] Abbas Razaghpanah, Arian Akhavan Niaki, Narseo Vallina-Rodriguez, Srikanth Sundaresan, Johanna Amann, and Phillipa Gill. 2017. Studying TLS usage in An-droid apps. In Conference on emerging Networking EXperiments and Technologies (CoNEXT) .[48] Abbas Razaghpanah, Rishab Nithyanand, Narseo Vallina-Rodriguez, Srikanth Sundaresan, Mark Allman, Christian Kreibich, Phillipa Gill, et al. 2018. Apps, trackers, privacy, and regulators: A global study of the mobile tracking ecosystem. In Network and Distributed System Security Symposium (NDSS) .[49] Joel Reardon, Kenneth A. Bamberger, and Serge Egelman. 2024. Anonymity, Consent, and Other Noble Lies: An Empirical Study of the Data Econ-omy. https://ibl.law.uiowa.edu/anonymity-consent-and-other-noble-lies-empirical-study-data-economy. Accessed: 2026-05-31. [50] Joel Reardon, Álvaro Feal, Primal Wijesekera, Amit Elazari Bar On, Narseo Vallina-Rodriguez, and Serge Egelman. 2019. 50 ways to leak your data: An exploration of apps’ circumvention of the android permissions system. In Proceedings of the USENIX Security Symposium .[51] Reuters. 2025. OpenAI Projected at Least 220 Million People Will Pay for ChatGPT by 2030, The Information Reports. Reuters. https://www.reuters.com/technology/openai-projected-least-220-million-people-will-pay-chatgpt-by-2030-information-2025-11-26/ Accessed: 2026-05-31. [52] Reuters. 2026. OpenAI to introduce ads to all ChatGPT free and Go users in US. https://www.reuters.com/business/media-telecom/openai-expand-ads-chatgpt-all-free-low-cost-users-information-reports-2026-03-21/ [53] Irwin Reyes, Primal Wijesekera, Joel Reardon, Amit Elazari Bar On, Abbas Raza-ghpanah, Narseo Vallina-Rodriguez, Serge Egelman, et al. 2018. “Won’t somebody think of the children?” examining COPPA compliance at scale. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[54] Franziska Roesner and David Kohlbrenner. 2026. Agentic Browsers and the Same-Origin Policy. In International Conference on Learning Representations (ICLR) .[55] Jaechul Roh, Eugene Bagdasarian, Hamed Haddadi, and Ali Shahin Shamsabadi. 2026. SPILLage: Agentic Oversharing on the Web. arXiv preprint arXiv:2602.13516

(2026). [56] Avishag Shapira, Parth Atulbhai Gandhi, Edan Habler, and Asaf Shabtai. 2025. Mind the web: The security of web use agents. arXiv preprint arXiv:2506.07153

(2025). [57] Singular. [n. d.]. Android SDK: Setting a User ID. Singular Developer Documen-tation. https://support.singular.net/hc/en-us/articles/35636052267803-Android-SDK-Setting-a-User-ID Accessed: 2026-05-31. [58] Thinkst Applied Research. 2026. Dockerized Canarytokens. https://github.com/ thinkst/canarytokens-docker. Accessed: 2026-05-27. [59] uBlock Origin Team. 2026. uBlock Origin. https://ublockorigin.com Accessed: 2026-05-24. [60] Alisha Ukani, Hamed Haddadi, Ali Shahin Shamsabadi, and Peter Snyder. 2025. Privacy Practices of Browser Agents. arXiv preprint arXiv:2512.07725 (2025). [61] Yash Vekaria, Aurelio Loris Canino, Jonathan Levitsky, Alex Ciechonski, Patricia Callejo, Anna Maria Mandalari, and Zubair Shafiq. 2025. Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative {AI } Assistants. In Proceedings of the USENIX Security Symposium .[62] Tim Vlummens, Aniketh Girish, Nipuna Weerasekara, Frederik Zuiderveen Bor-gesius, Gunes Acar, and Narseo Vallina-Rodriguez. 2026. Bridges to Self: Silent Web-to-App Tracking on Mobile via Localhost. In Proceedings of the USENIX Security Symposium .[63] Ethan Wang, Zubair Shafiq, and Yash Vekaria. 2026. FP-Agent: Fingerprinting AI Browsing Agents. arXiv preprint arXiv:2605.01247 (2026). [64] Nipuna Weerasekara, José Miguel Moreno, Srdjan Matic, Joel Reardon, Juan Tapi-ador, Narseo Vallina-Rodríguez, et al. 2025. Tracking without borders: Studying the role of webviews in bridging mobile and web tracking. In Proceedings on Privacy Enhancing Technologies (PoPETs) .[65] Wired. 2026. OpenAI Enables Cookies by Default for Free ChatGPT Users. https://www.wired.com/story/openai-enables-cookies-by-default-for-free-chatgpt-users/ Accessed: 2026-05-31. [66] Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel, and Chaowei Xiao. 2024. A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems. arXiv:2402.18649 [cs.CR] https://arxiv.org/abs/2402.18649 [67] Yuhao Wu, Evin Jaff, Ke Yang, Ning Zhang, and Umar Iqbal. 2025. An in-depth investigation of data collection in llm app ecosystems. In Proceedings of the Internet Measurement Conference (IMC) .[68] Yuhao Wu, Ke Yang, Franziska Roesner, Tadayoshi Kohno, Ning Zhang, and Umar Iqbal. 2025. Towards Automating Data Access Permissions in AI Agents. arXiv:2511.17959 [cs.CR] https://arxiv.org/abs/2511.17959 [69] Zhonghao Zhan, Huichi Zhou, Zhenhao Li, Peiyuan Jing, Krinos Li, and Hamed Haddadi. 2026. How Adversarial Environments Mislead Agentic AI? arXiv preprint arXiv:2604.18874 (2026). 15 Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al.

A Conversation Titles

This section provides illustrative examples of how different AI ser-vices automatically synthesize user intent to generate chat session titles. As shown in Table 7, the web implementations of ChatGPT, Claude, and Grok result in different outcomes regarding brevity, cap-italization, and information density when condensing user prompts.

Table 7: Examples of automatically generated conversation titles derived from concise user prompts in the web imple-mentations of ChatGPT, Claude and Grok.

Prompt LLM Generated Title

What are the symptoms of early-stage Parkinson’s disease?

ChatGPT Early-stage Parkinson’s Symptoms Claude Early-stage Parkinson’s dis-ease symptoms Grok Early Parkinson’s Disease Symptoms

My salary is $85k. How much mortgage can I afford in NYC?

ChatGPT Mortgage Affordability in NYC Claude Mortgage affordability on $85k salary in NYC Grok $85k NYC Salary: $280k-$350k Mortgage

B CSP Relationships

We detail the observed CSP relationships identified across the ana-lyzed ecosystem. Specifically, Table 8 shows the number of different conversational AI providers that integrate these trackers into their CSP policies, highlighting the most prevalent third-party advertis-ing and analytics domains.

Table 8: Observed CSP relationships.

Organization Domain(s) # Providers Google Tag Manager googletagmanager.com 6Google Analytics google-analytics.com 5Google Ads googleadservices.com 4Google DoubleClick doubleclick.net 3Meta Pixel / Conversions API connect.facebook.net, 4facebook.net TikTok Analytics analytics.tiktok.com 2Reddit Ads / Analytics pixel-config.reddit.com, 2redditstatic.com Microsoft Bing bat.bing.com, bing.com 2Intercom intercom.io, intercomcdn.com 2

C Consent Implementations

To provide context on Privacy policies and the differences in im-plementation between providers, Figures 4, 5, 6 show snapshots of what cookie consent banners looked like at the time of capture.

Figure 4: Mistral web consent banner.

Figure 5: ChatGPT web consent banner.

Figure 6: Gemini web consent banner.

D Sample data sent

We list representative samples of data transmitted during interac-tions in Figures 7, 8, 9, 10, 11. The examples illustrate the informa-tion observed in requests sent to analytics, telemetry, and advertis-ing endpoints, including conversation identifiers, titles, prompts, URLs, user identifiers, and tracking pixels. Sensitive values have been redacted for readability and privacy.

16

Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents Proceedings on Privacy Enhancing Technologies YYYY(X)

TikTok Pixel --- POST analytics . tiktok . com / api / v2 / pixel / act " user ":{" anonymous_id ": "< _ttp >", ...} , " page ":{ " url ": " https :// grok . com /c/< convUUID >? rid =..." , ... }, " meta ":{" title ": "< convTitle >", ...} Meta Pixel PageView --- GET www . facebook . com / tr /?... ev = PageView dl = https :// grok . com /c/< convUUID >? rid =... fbp = < _fbp > pmd [ title ] = < convTitle > Meta Pixel SubscribedButtonClick ( share ) --- GET www . facebook . com / tr /?... ev = SubscribedButtonClick dl = https :// grok . com /c/< convUUID >? rid =... cd [ buttonText ] = Share cd [ pageFeatures ] = {" title ":" < convTitle > "} fbp = < _fbp > sGTM relay --- POST sgtm - prod -985009374134. us - central1 . run . app / data ? event = sent_3_chat_messages " page_location ": " https :// grok . com /c/< convUUID >? rid =..." , " page_referrer ": " https :// accounts .x. ai /" , " page_title ": "< convTitle >", " common_cookie ": {" _fbp ": "< _fbp >", " _ttp ": "< _ttp >" }, " _dcid_temp ": "< _dcid >" Twitter Pixel --- GET analytics . twitter . com /1/ i/ adsct ?... pt = < convTitle > tw_document_href = https :// grok . com /c/< convUUID >? rid =... twpid = < _twpid > Google Ads --- GET www . googleadservices . com / pagead / conversion /... url = https :// grok . com /c/< convUUID >? rid =... ref = https :// accounts .x. ai / tiba = < convTitle > em = tv .1~ em .< userHash >

Figure 7: A single Grok conversation propagates to seven advertising and analytics trackers. Every recipient receives the same conversation identifier and the automatically gen-erated title. Meta ( _fbp ), TikTok ( _ttp ), and Twitter ( _twpid )cookies are synced across the flows—the Meta and TikTok cookies also travel server-to-server through the sGTM re-lay, invisible to ad blockers—and Google Ads additionally receives hashed user information ( em ).

GET region1 . google - analytics . com /g/ collect ?... cid = < cid > dl = https :// gemini . google - b197145817 . com / app / dt = < convTitle >

Figure 8: Google Analytics collects conversation Title from Gemini.

POST browser - intake - us5 - datadoghq . com / api / v2 / spans " resource ": "/ organizations /{ organization }/ chat_conversations /{ chat }/ share " " http . url ": " https :// claude . ai / api / organizations /< orgUUID >/ chat_conversations /< convUUID >/ share ", " usr ": {" id ": "< userUUID >", " account_uuid ": "< userUUID >", " organization_id ": "< orgUUID >", " subscription_level ": " free " }, " session ": { " id ": "< sessionUUID >" }, " view ": { " id ": "< viewUUID >" }, " user_action ": { " id ": "< actionUUID >" }, " device ": { " model ": " AOSP on sargo ", " ram_mb ": 3577 }, " os ": { " name ": " Android ", " version ": "12" }

Figure 9: Datadog span emitted by the Claude mobile app for a conversation-share action. The Anthropic organization, conversation, and share UUIDs travel in the resource URL alongside account identifiers; further session, view, and per-action UUIDs let the recipient stitch every gesture into the same authenticated profile. The share resource is created without explicit user interaction.

POST browser - intake - us5 - datadoghq . com / api / v2 / rum ?dd - api -key =... {" type ": " long_task ", " application ":{" id ":" < applicationUUID > "} , " view ": {" url ":" https :// claude . ai / share /{ id }" , ...} , " session ":{" id ":" < sessionUUID > " ," type ":" user "} , " usr ": {" anonymous_id ":" < anonUUID > " ,...} , " long_task ":{ " scripts ":[{ " source_url ":" https :// claude . ai / share /< shareUUID >", " invoker ": " https :// claude . ai / share /< shareUUID >" }] }, ... }

Figure 10: Datadog RUM beacon emitted by the Claude web

client. The shared-conversation permalink is reported as the long-task script source and invoker, alongside application and session identifiers.

17 Proceedings on Privacy Enhancing Technologies YYYY(X) Oliveira et al.

TikTok Pixel ---POST analytics . tiktok . com / api / v2 / pixel / act " page ":{ " url ": " https :// grok . com / share /< shareUUID >? rid =..." , ... }, " open_graph ":{ " og : image ": " https :// grok . com / share /< shareUUID >/ opengraph - image /< shareUUID >", " og : image : alt ": "< convScreenshotPreview >", ... }, " meta ":{ " title ": "< convTitle >", " meta : description ": "< UserPrompt >" }Meta Pixel PageView ---GET www . facebook . com / tr /?... ev =PageView dl =https :// grok . com / share /< shareUUID >? rid =... fbp =< _fbp > pmd [ title ] =< convTitle > pmd [ description ] =< UserPrompt >

Figure 11: A shared Grok conversation propagates extra infor-mation about the conversation to Meta and TikTok advertis-ing and analytics trackers. Both recipients receives the same shared conversation identifier, the automatically generated title and the last user prompt. Meta ( _fbp ) and TikTok ( _ttp )cookies are synced across the flows. Additionally, TikTok receives a screenshot of the most recent part of the conversa-tion, as shown in Figure 12.

E Conversation Screenshots Leakage

We captured Grok’s screenshot as transmitted to TikTok when a shared conversation is accessed, as shown in Figure 11. The screen-shot captures a leaked user conversation in Grok, as it is illustrated in Figure 12.

Figure 12: An example of the exact screenshot that gets sent to TikTok by Grok on shared interactions.

来源:Hacker News:AI 热帖 · jorgegarciaherrero.com