Gemini models are Google’s Transformer-based multimodal AI models designed to process and generate across modalities such as text, images, audio, video, and code. They are used in Google products and developer tools, including agentic and conversational applications.
Google's state-of-the-art Transformer-based multimodal models. Gemini models are specifically designed to integrate with agents. Users can interact with Gemini models in a variety of ways, including through an interactive dialog interface and through SDKs.