Back to feed
arXiv cs.CL·

Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation

Signal
75
Hype
25
In three linesAgentViSS benchmark evaluates visual social intelligence of multimodal agents in social simulations. 240 scenarios, 585 roles, 2,340 instances test whether MLLMs use visual cues (expressions, posture, gaze) to guide interactions. Seven models evaluated show gap: expression and conflict handling near saturation, interaction regulation and visually grounded outcomes remain substantially harder.
Read source
Your take?
BenchmarksVisionAI AgentsMulti-agent

Summary generated by Claude — human-verified