newsfilter.io
Fireside Chat, Interview, Conference Presentation

Fast, Efficient Inference at Scale with Heterogeneous Hardware | Gimlet Labs x d-Matrix | RAISE 2026

  • Company Overview & Funding

    • Zain Asghar is the founder and CEO of Gimlet Labs, which has raised nearly $80 million in a Series A round led by Menlo Ventures.
    • The company is building the industry's first multi-silicon inferencing cloud.
    • Prior to Gimlet, Asghar founded Pixie Labs (acquired by New Relic), served as a professor at Stanford, an architect at NVIDIA, and an engineering lead at Google.
  • Core Technology & Value Proposition

    • Gimlet aims to make AI inference workloads 10 times more efficient by decomposing them into fine-grained components (including individual layers) and distributing them across heterogeneous hardware.
    • The platform moves beyond simple multi-vendor compatibility to true multi-silicon orchestration, allocating specific workloads to CPUs, GPUs, and NPUs based on their architectural strengths.
    • The system is designed to handle the diversity of agentic workloads, which involve tool calls and mixed CPU/GPU tasks.
    • By balancing memory-bound vs. compute-bound workloads across different silicon types, Gimlet can increase capacity by 4x to 5x at the same cost and improve user-perceived interactivity by a similar margin.
  • Market Trends & Strategic Shifts

    • The push toward multi-silicon is driven by scarcity of resources in data centers and the increasing complexity of agentic AI workloads.
    • The industry is shifting from a "zero-sum" GPU vs. alternative chip mindset to an "augmentation" model where different chips optimize specific parts of the workload.
    • Gimlet views GPUs as the "CPUs of the AI era" but asserts that other silicon types are necessary to accelerate specific segments for better overall efficiency.
  • Key Applications

    • Primary use cases target coding agents, where order-of-magnitude speed improvements are critical for developer workflows.
    • The team is currently running proprietary coding agents and models internally on their own stack to validate performance and experience.
  • Hardware Partnership Strategy

    • Gimlet partners with chipmakers based on specific architectural trade-offs between compute density, memory bandwidth, and memory capacity.
    • D-Matrix: Selected as a key partner for its stacked memory architecture, which addresses the bandwidth and capacity limitations of SRAM-centric chips by placing memory directly on top of compute.
    • Target Hardware: The platform supports LPX, Cerebrus, and other NPUs to utilize the specific strengths of each chip.
    • Partnership Criteria: Gimlet evaluates partners on how well they fit into a full-stack solution that requires balancing high memory bandwidth with compute efficiency.
  • Technical Challenges & Full-Stack Approach

    • The company has become vertically integrated (operating its own data centers) because no external entity currently deploys true heterogeneous hardware clusters.
    • Key challenges in heterogeneous orchestration include:
      • Software: Incompatible programming models, particularly with spatial architectures where GPU programming paradigms do not translate.
      • Networking: Lack of standard interconnects (e.g., NVLink) and high sensitivity to network latency during disaggregation.
      • Infrastructure: Incompatibilities in cooling, electrical systems, and rack topologies when mixing different chip types.
    • Gimlet is building the necessary network topologies and managing data center capacity to support these multi-silicon deployments.
  • Entrepreneurial Learnings & Vision

    • Asghar applies the lesson from his previous venture to avoid "middle-ground" decision-making, favoring high-conviction bets and late-binding to less critical decisions.
    • The company is actively redefining data center design by moving away from GPU-only scaling toward inference-centric systems utilizing new hardware types.
    • Asghar believes the industry is in the early innings of the inferencing era, with significant optimization and efficiency gains still to be realized.