
🔍MLLMs are nearsighted: zooming helps but kills speed. What if we teach them to "see" fine details without zooming? Introducing Region-to-Image Distillation(R2I): we train models to internalize "zooming". 🏆 ZwZ-8B achieves SOTA performance on fine-grained perception with zero tool calls. 📊 Plus ZoomBench: human-AI collaboration construction, dual-view "zooming gap" metric Dive into our joint efforts: 📑Paper: https://t.co/tdEn2QwoKD 🔗Code: https://t.co/b8S4HqcdPT ⚙️ Model&Data: https://t.co/5ZDak0L3C8 #OpenSource #MLLM #LLMs #inclusionAI
If a model internalizes zooming, a perception drops an entire tool-call loop and its latency. ZwZ-8B and ZoomBench provide a checkpoint and a to test that trade in your own pipeline.
Checking sign-in…
Loading comments…