newsfilter.io
Conference Presentation, Keynote, Product Demonstration

Where We're Going We Don't Need Keyboards, Just Voice | Neil Zeghidour, Gradium | RAISE 2026

  • Company & Mission Profile

    • Gradium is a spin-off of the non-profit research lab Qtize, recently listed among "foundation frontier labs."
    • The company focuses on training voice models for real-time interaction, offering capabilities in transcription, synthesis, translation, and transformation.
    • Gradium aims to commercialize frontier research to power the "best model to power voice AI."
    • Qtize previously introduced the first speech-to-speech conversational models, real-time speech translation systems, and on-device CPU models.
  • Shift in Interaction Paradigm

    • For four decades, computer interaction has relied on keyboards and structured machine language (e.g., URLs, code syntax) rather than natural human language.
    • Programming interaction has evolved from unintelligible binary (punch cards) to assembly, then high-level languages, and now to conversational interaction with coding agents.
    • Voice is identified as the most natural medium for natural language, allowing users to speak three times faster than they type.
    • The transition reflects a move where machines now understand the subtleties of human language, reversing the historical need for humans to adapt to machine syntax.
  • Technological Evolution of Voice AI

    • Closed-Ended Agents (Pre-LLM): Systems like the 2011 Siri demo used complex logic to map speech to specific intents within constrained domains (weather, stocks, contacts) without open conversation.
    • Conversational LLMs: Early voice modes utilized LLMs to enable open conversation but lacked agentic capabilities (tool calling) and suffered from high latency (5–6 seconds).
    • Voice Agents: Current systems combine conversational LLMs with tool calling and database access to perform agentic tasks, such as ordering in a fast-food drive-thru with specific dietary restrictions.
    • Full-Duplex Models: New architectures (e.g., Gradium's "Moshi," OpenAI's recent bidirectional model) utilize a single audio-to-audio model to eliminate turn-taking latency.
      • These systems operate in continuous mode, allowing users to speak over the AI without breaking the conversation flow.
      • This contrasts with cascaded systems (STT + LLM + TTS) that fail during interruptions or background noise.
  • Technical Challenges and Next Frontiers

    • Robotics Environment: The next frontier involves deploying voice AI in noisy, multi-microphone settings (e.g., factory robots) rather than clean, single-user phone calls.
    • Acoustic Scene Analysis: Robots must filter background noise and irrelevant speech from multiple simultaneous speakers.
    • Diarization: Systems require the ability to distinguish "who said what" in multi-party conversations to resolve conflicting instructions.
    • Long-Form Interaction: Industrial applications demand models capable of maintaining context during very long sessions throughout a workday.
  • Market Applications & Trends

    • Customer Support: High-growth sector utilizing infinite-patience agents for scheduling and troubleshooting.
    • Gaming: Interactive NPCs, AI coaches, and real-time commentary systems.
    • Education: Language learning platforms focusing on conversational practice.
    • Software Development: Shift toward voice-based management of sub-agents for coding and task delegation.
    • Strategic Vision: Voice is projected to become the primary operating system for machines, characterized by always-on, asynchronous, and hierarchical agent interaction.
  • Business & Operational Updates

    • Gradium is offering free access to its models via Gradium.ai.
    • The company is actively recruiting technical and GTM (Go-To-Market) talent.
    • Gradium is based in Paris and has announced funding to open a new office in San Francisco.