Fireside Chat, Interview, Conference Presentation
Fast, Efficient Inference at Scale with Heterogeneous Hardware | Gimlet Labs x d-Matrix | RAISE 2026
Company Overview & Funding
- Zain Asghar is the founder and CEO of Gimlet Labs, which has raised nearly $80 million in a Series A round led by Menlo Ventures.
- The company is building the industry's first multi-silicon inferencing cloud.
- Prior to Gimlet, Asghar founded Pixie Labs (acquired by New Relic), served as a professor at Stanford, an architect at NVIDIA, and an engineering lead at Google.
Core Technology & Value Proposition
- Gimlet aims to make AI inference workloads 10 times more efficient by decomposing them into fine-grained components (including individual layers) and distributing them across heterogeneous hardware.
- The platform moves beyond simple multi-vendor compatibility to true multi-silicon orchestration, allocating specific workloads to CPUs, GPUs, and NPUs based on their architectural strengths.
- The system is designed to handle the diversity of agentic workloads, which involve tool calls and mixed CPU/GPU tasks.
- By balancing memory-bound vs. compute-bound workloads across different silicon types, Gimlet can increase capacity by 4x to 5x at the same cost and improve user-perceived interactivity by a similar margin.
Market Trends & Strategic Shifts
- The push toward multi-silicon is driven by scarcity of resources in data centers and the increasing complexity of agentic AI workloads.
- The industry is shifting from a "zero-sum" GPU vs. alternative chip mindset to an "augmentation" model where different chips optimize specific parts of the workload.
- Gimlet views GPUs as the "CPUs of the AI era" but asserts that other silicon types are necessary to accelerate specific segments for better overall efficiency.
Key Applications
- Primary use cases target coding agents, where order-of-magnitude speed improvements are critical for developer workflows.
- The team is currently running proprietary coding agents and models internally on their own stack to validate performance and experience.
Hardware Partnership Strategy
- Gimlet partners with chipmakers based on specific architectural trade-offs between compute density, memory bandwidth, and memory capacity.
- D-Matrix: Selected as a key partner for its stacked memory architecture, which addresses the bandwidth and capacity limitations of SRAM-centric chips by placing memory directly on top of compute.
- Target Hardware: The platform supports LPX, Cerebrus, and other NPUs to utilize the specific strengths of each chip.
- Partnership Criteria: Gimlet evaluates partners on how well they fit into a full-stack solution that requires balancing high memory bandwidth with compute efficiency.
Technical Challenges & Full-Stack Approach
- The company has become vertically integrated (operating its own data centers) because no external entity currently deploys true heterogeneous hardware clusters.
- Key challenges in heterogeneous orchestration include:
- Software: Incompatible programming models, particularly with spatial architectures where GPU programming paradigms do not translate.
- Networking: Lack of standard interconnects (e.g., NVLink) and high sensitivity to network latency during disaggregation.
- Infrastructure: Incompatibilities in cooling, electrical systems, and rack topologies when mixing different chip types.
- Gimlet is building the necessary network topologies and managing data center capacity to support these multi-silicon deployments.
Entrepreneurial Learnings & Vision
- Asghar applies the lesson from his previous venture to avoid "middle-ground" decision-making, favoring high-conviction bets and late-binding to less critical decisions.
- The company is actively redefining data center design by moving away from GPU-only scaling toward inference-centric systems utilizing new hardware types.
- Asghar believes the industry is in the early innings of the inferencing era, with significant optimization and efficiency gains still to be realized.