Conference Presentation, Keynote, Product Demonstration
What's New for Startups with Google DeepMind? | Omar Sanseviero | RAISE Summit 2026
Organizational Scope & Mission
- Omar Sanseviero leads Developer Experience at Google DeepMind, overseeing all model launches and API integrations.
- Core responsibilities include managing the Gemini ecosystem, Google AI Studio, Gemma open models, and various GenMedia initiatives.
- The strategic goal is to provide bold, responsible AI integration that enables rapid development from "prompt to production" for developers globally.
Core Model Family: Gemini
- Gemini Family Overview:
- Gemini is established as Google's most powerful model foundation, providing world understanding, physics simulation, and multimodal reasoning.
- The family offers a spectrum of sizes to balance latency, cost, and complexity trade-offs.
- Specific Model Releases & Roadmap:
- Gemini 3.5 Flash: Released two months ago; optimized for speed and cost-efficiency.
- Gemini 2.5 Pro: Scheduled for imminent release.
- Gemini 3.1 Flash: A text-to-speech model featuring "director style control" for tone, accent, and persona management.
- Gemini 3.5 Flash 2.0: Mentioned as an upcoming iteration.
- Advanced Capabilities:
- Supports long-horizon tasks requiring complex planning, agent orchestration, and multi-day execution times.
- Enables agentic workflows where models can perform Google Search or function calls to ground generation in real-world data (e.g., weather, specific biographical details).
- Gemini Family Overview:
Generative Media (GenMedia) Models
- NanoBanana (Image Generation):
- Built on the Gemini foundation to add image generation and conversational image editing capabilities.
- NanoBanana 2 Lite: Released one week prior to the talk; significantly faster, cheaper, and higher-performing than version 1.
- Performance metric: Generates a 3x3 grid with complex camera angle variations in approximately 3–4 seconds.
- Supports tool use (e.g., searching for image context) and iterative conversation-based editing.
- Vio (Video & Audio):
- A text-to-video model capable of generating synchronized audio alongside visual content.
- Gemini OVNI (Anything-to-Anything):
- Initial Release (Gemini OVNI Flash): Supports text, audio, and video inputs to generate high-quality video.
- Allows for video references as inputs and real-time conversation about generated video content.
- Designed to incorporate world understanding to maintain consistency across modalities.
- Future iterations expected within the next two months.
- NanoBanana (Image Generation):
Live & Real-Time APIs
- Live API:
- Provides ultra-low latency voice and vision understanding for real-time interaction.
- Capable of processing live screen shares or video feeds to provide contextual guidance via speech and text.
- Supports fluid, interruptible conversations across multiple languages (English, Spanish, French).
- Live Translate API:
- Released two weeks prior to the talk; enables real-time audio-to-audio translation.
- Allows simultaneous speech in one language (e.g., English) with generation in another (e.g., Hindi, Portuguese) without latency penalties.
- Live API:
Open Models (Gemma)
- Gemma 4 Release:
- Latest iteration released three months prior; available under the Apache 2.0 license.
- Architecture: Ranges from 2 billion to 32 billion parameters; supports multimodal inputs (text, image, audio, video).
- Context Window: 256,000 tokens, balancing high capacity with consumer-grade hardware feasibility.
- Multilingual Support: Trained on and performs well in 140 languages, with strong performance in European, Korean, and Japanese languages.
- Capabilities: Includes reasoning, instruction following, and function calling.
- Deployment & Adoption:
- Designed to run on consumer hardware, including gaming GPUs and mobile devices (Pixel, iPhone).
- Achieved over 250 million downloads within the first three months.
- Demonstrated offline capabilities on mobile phones and AR glasses for tasks including OCR, math reasoning, and visual identification.
- Gemma 4 Release:
Developer Tooling & Infrastructure
- Google AI Studio:
- Free interactive playground for testing, prototyping, and deploying applications.
- Facilitates direct app creation and production deployment with minimal friction.
- Gemini API Managed Agents:
- Allows developers to fully manage agent lifecycles within Google infrastructure without managing sandboxes or security setups.
- Google AI Gravity:
- Agentic ID and harness for streamlined agentic development.
- Ecosystem Integration:
- Prioritizes compatibility with existing open-source tools (e.g., LiveKit, diverse SDKs) to minimize switching costs.
- Google AI Studio:
Future Applications & Vertical Use Cases
- Robotics:
- Application of Gemini for embodied reasoning and vision-action tasks.
- Specific capabilities include precise movement and pointing via vision models.
- Genie 3:
- A world generation model (not API-accessible) that creates interactive, physics-based environments similar to a game engine.
- Weather Next:
- An open model used globally for weather forecasting and deployed in production by research institutes.
- Fire Detection:
- Improved sensitivity detects fires as small as one or two classrooms in satellite imagery, previously requiring fires the size of football stadiums.
- Enables early response teams to mitigate fires more effectively.
- Broader Vision:
- Strategic focus on intersections with healthcare, finance, bio, and robotics.
- Goal is to provide a foundational layer for startups and enterprises to build specialized applications.
- Robotics: