Google DeepMind has launched Gemini Omni 1.1 Flash, an incremental update to its multimodal model that introduces developer controls for tuning output behavior across text, image, audio, and video modalities.
The update focuses on giving production teams more predictable and controllable outputs — a common pain point for enterprises deploying multimodal AI at scale. Developers can now adjust parameters for output format, style consistency, and modality-specific behavior without fine-tuning.
Gemini Omni 1.1 Flash is designed for applications that need to combine multiple output types in a single workflow, such as customer-service agents that generate text responses alongside visual aids, or content-creation tools that produce coordinated text-and-image outputs.
The model is available through the Gemini API and Google Cloud, with pricing positioned for high-throughput production workloads.