这不只是建模问题。也是基准测试问题。
当前多模态模型靠语言捷径'作弊',真实场景落地将暴露致命隐患
this isn’t just a modeling problem. it’s also a benchmarking problem.
spurious correlations are always a pain, but in multimodal llms they become a particularly tough battle. On one hand, you want to leverage the language prior to enable better generalization; on the other, that same language prior can turn into a shortcut that makes the model effectively blind.
the irony is that humans do the same thing. We still gravitate toward language-first tasks, and the “multimodal results” in major model releases like gpt-5 reflect exactly that bias.
I mean, economically this makes most sense for LLM companies: you can claim wins in “multimodal reasoning” without investing heavily in real multimodal research.
that shortcut will come due tho. when you try to put these systems into glasses, robots, or anything else that touches the real world, the cracks will show. and they’ll be costly.
I couldn’t believe GPT-5 could make this mistake until @ziqiao_ma pointed it out to me. Highly recommend this paper (https://arxiv.org/abs/2406.16860) on vision-centric evaluation of multimodal LLMs from @sainingxie — now imagine the same rigor applied to VLAs.在 X 查看被引用的帖子
来源:Saining Xie · x.com