Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Source
deepmind.google
Date
Terms in this piece · Glossary
multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
Why it matters
An encoder-free multimodalA model that works with more than text — reading images, audio, or video, and sometimes generating them too.Full definition → open model small enough to run on a laptop, relevant if you are moving vision-and-text workloads onto local hardware instead of an API.