原文英文,约400词,阅读约需2分钟。
📝
内容提要
Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets...