newsfilter.io
Fireside Chat, Interview, Conference Presentation

Fast, Efficient Inference at Scale with Heterogeneous Hardware | Gimlet Labs x d-Matrix | RAISE 2026

  • Aims to achieve approximately 10x efficiency gains for inference workloads and deliver 4x to 5x improvements in both capacity per cost and user interactive performance.
  • Plans to partition workloads into fine-grained components—such as individual models, layers, or granular units—and distribute them across heterogeneous hardware including CPUs, GPUs, and specialized accelerators like D-Matrix, Cerebras, and LPX.
  • Targets coding agents operating with human-in-the-loop workflows to enable instant experiences that allow engineers to match their thinking speed.
  • Intends to scale solutions by increasing available chip types and improving efficiency to create a "double effect" that generates greater capacity while addressing current data center resource scarcity.
  • Plans to deploy a small, Gimlet-orchestrated installation of multi-silicon hardware internally for proprietary coding agents and models.
  • Leverages trade-offs between compute and memory bandwidth to balance memory-bound versus compute-bound workloads, with D-Matrix specifically targeting stacked memory to maximize available area.
  • Expects D-Matrix to eventually provide chips with unlimited capacity and bandwidth, though current iterations face networking compatibility issues as NVLink cannot effectively cross different chip types.
  • Views the challenge as a full-stack issue extending from software programming models and spatial architectures through to cooling, electrical systems, and data center topology design.
  • Anticipates that new software stacks for neo-accelerators require distinct approaches compared to past GPU programming models due to spatial architecture differences.
  • Plans to build and manage significant production data center capacity, redefining future construction to handle sensitivity to network latency and support true disaggregation.
  • Shifts industry mindset from a zero-sum competition against GPUs to an approach of augmenting GPU capabilities with diverse chip types to accelerate specific workload pieces.
  • Believes GPUs function as the AI era's equivalent of CPUs but expects significant further optimizations and efficiency gains as the inferencing era remains in its early innings.
  • Expects agentic workloads to increase the diversity of AI inference through expanded tool calls and CPU operations.
  • Identifies deployment risks regarding data center cooling and electrical systems, alongside potential challenges in building topologies that manage network and latency sensitivities.