Qwen2.5-Omni is an openly available 7B end-to-end model that perceives text, images, audio, and video while both text and natural speech responses in real time — useful if you want to build voice/vision agents on self-hostable weights rather than closed APIs.
“We release Qwen2.5-Omni, the new flagship end-to-end multimodal model in the Qwen series.”
Checking sign-in…
Loading comments…