Inverse Graphics as RL Environments: Testing Whether VLMs Can Actually See
Vision-language models can describe a scene in paragraph-length detail and still fail to tell you whether a red cube is in front of or behind a blue cylinder.
1 article tagged with #spatial-reasoning.
Vision-language models can describe a scene in paragraph-length detail and still fail to tell you whether a red cube is in front of or behind a blue cylinder.