Vibeleaderboard
← All Intel
Intel / article

MLLMs Fail to Refuse when Using Tools Agentically

Source
arxiv.org
Author
Rikiya Takehi, Ryo Hachiuma, Shaona Ghosh, Dan Zhao, Yu-Chiang Frank Wang, Yusuke Hirota
Date
Why it matters

Safety behavior measured without tools overstates safety once an can call tools. Engineers building tool-using agents should not assume refusal behavior carries over and should evaluate safety in the agentic setting.

Terms in this piece · Glossary
  • multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Recommended reads
Comments

Checking sign-in…

Loading comments…