newsfilter.io
Conference Presentation, Keynote, Product Demonstration

What's New for Startups with Google DeepMind? | Omar Sanseviero | RAISE Summit 2026

  • Organizational Scope & Mission

    • Omar Sanseviero leads Developer Experience at Google DeepMind, overseeing all model launches and API integrations.
    • Core responsibilities include managing the Gemini ecosystem, Google AI Studio, Gemma open models, and various GenMedia initiatives.
    • The strategic goal is to provide bold, responsible AI integration that enables rapid development from "prompt to production" for developers globally.
  • Core Model Family: Gemini

    • Gemini Family Overview:
      • Gemini is established as Google's most powerful model foundation, providing world understanding, physics simulation, and multimodal reasoning.
      • The family offers a spectrum of sizes to balance latency, cost, and complexity trade-offs.
    • Specific Model Releases & Roadmap:
      • Gemini 3.5 Flash: Released two months ago; optimized for speed and cost-efficiency.
      • Gemini 2.5 Pro: Scheduled for imminent release.
      • Gemini 3.1 Flash: A text-to-speech model featuring "director style control" for tone, accent, and persona management.
      • Gemini 3.5 Flash 2.0: Mentioned as an upcoming iteration.
    • Advanced Capabilities:
      • Supports long-horizon tasks requiring complex planning, agent orchestration, and multi-day execution times.
      • Enables agentic workflows where models can perform Google Search or function calls to ground generation in real-world data (e.g., weather, specific biographical details).
  • Generative Media (GenMedia) Models

    • NanoBanana (Image Generation):
      • Built on the Gemini foundation to add image generation and conversational image editing capabilities.
      • NanoBanana 2 Lite: Released one week prior to the talk; significantly faster, cheaper, and higher-performing than version 1.
      • Performance metric: Generates a 3x3 grid with complex camera angle variations in approximately 3–4 seconds.
      • Supports tool use (e.g., searching for image context) and iterative conversation-based editing.
    • Vio (Video & Audio):
      • A text-to-video model capable of generating synchronized audio alongside visual content.
    • Gemini OVNI (Anything-to-Anything):
      • Initial Release (Gemini OVNI Flash): Supports text, audio, and video inputs to generate high-quality video.
      • Allows for video references as inputs and real-time conversation about generated video content.
      • Designed to incorporate world understanding to maintain consistency across modalities.
      • Future iterations expected within the next two months.
  • Live & Real-Time APIs

    • Live API:
      • Provides ultra-low latency voice and vision understanding for real-time interaction.
      • Capable of processing live screen shares or video feeds to provide contextual guidance via speech and text.
      • Supports fluid, interruptible conversations across multiple languages (English, Spanish, French).
    • Live Translate API:
      • Released two weeks prior to the talk; enables real-time audio-to-audio translation.
      • Allows simultaneous speech in one language (e.g., English) with generation in another (e.g., Hindi, Portuguese) without latency penalties.
  • Open Models (Gemma)

    • Gemma 4 Release:
      • Latest iteration released three months prior; available under the Apache 2.0 license.
      • Architecture: Ranges from 2 billion to 32 billion parameters; supports multimodal inputs (text, image, audio, video).
      • Context Window: 256,000 tokens, balancing high capacity with consumer-grade hardware feasibility.
      • Multilingual Support: Trained on and performs well in 140 languages, with strong performance in European, Korean, and Japanese languages.
      • Capabilities: Includes reasoning, instruction following, and function calling.
    • Deployment & Adoption:
      • Designed to run on consumer hardware, including gaming GPUs and mobile devices (Pixel, iPhone).
      • Achieved over 250 million downloads within the first three months.
      • Demonstrated offline capabilities on mobile phones and AR glasses for tasks including OCR, math reasoning, and visual identification.
  • Developer Tooling & Infrastructure

    • Google AI Studio:
      • Free interactive playground for testing, prototyping, and deploying applications.
      • Facilitates direct app creation and production deployment with minimal friction.
    • Gemini API Managed Agents:
      • Allows developers to fully manage agent lifecycles within Google infrastructure without managing sandboxes or security setups.
    • Google AI Gravity:
      • Agentic ID and harness for streamlined agentic development.
    • Ecosystem Integration:
      • Prioritizes compatibility with existing open-source tools (e.g., LiveKit, diverse SDKs) to minimize switching costs.
  • Future Applications & Vertical Use Cases

    • Robotics:
      • Application of Gemini for embodied reasoning and vision-action tasks.
      • Specific capabilities include precise movement and pointing via vision models.
    • Genie 3:
      • A world generation model (not API-accessible) that creates interactive, physics-based environments similar to a game engine.
    • Weather Next:
      • An open model used globally for weather forecasting and deployed in production by research institutes.
    • Fire Detection:
      • Improved sensitivity detects fires as small as one or two classrooms in satellite imagery, previously requiring fires the size of football stadiums.
      • Enables early response teams to mitigate fires more effectively.
    • Broader Vision:
      • Strategic focus on intersections with healthcare, finance, bio, and robotics.
      • Goal is to provide a foundational layer for startups and enterprises to build specialized applications.