Google DeepMind released Gemma 4 12B on June 4, the first open-source multimodal model with an encoder-free architecture. The "encoder-free" design skips the separate vision encoder, instead having the language model directly process image tokens. This simplifies the architecture, reduces latency, and enables more flexible multimodal interactions, with the model running on-device via Google AI Edge.