Latest Interviews
Showing 1–4 of 4 interview transcripts.
Clear all filters- RAISE Summit19 min
Inside the AI Inference Cluster: Measuring What Matters | Mansour Karam, Aria Networks | RAISE 2026
The AI infrastructure sector is pivoting from homogeneous general-purpose systems to specialized architectures that optimize distinct model training and serving economics, with high-throughput inference offering potential cost savings of up to 100x. This shift demands deep networking solutions capable of managing dynamic, unpredictable traffic patterns and microsecond-scale latency requirements for complex agent workflows. Leading this transition, ARIA proposes a multi-layered AI agent architecture that integrates full-stack telemetry to dynamically coordinate hardware, software, and network operations across immediate physical reactions and long-term strategic planning.
- RAISE Summit18 min
Fast by Design: The Infrastructure Powering the Agentic Era | Rodrigo Liang | RAISE Summit 2026
Rodrigo Liang, Monti Saroya, Dylan Patel
SambaNova has secured the first close of a $1 billion funding round at an $11 billion valuation, establishing itself as JP Morgan's preferred inference provider for multi-trillion parameter models. The company differentiates its heterogeneous architecture by delivering up to 10x faster decode performance in full precision, avoiding the accuracy trade-offs of quantization while integrating with NVIDIA and Intel infrastructure. By leveraging a distribution network of 94 subsidiaries, SambaNova targets enterprise adoption through hybrid air-gapped and cloud solutions that prioritize data sovereignty and model choice over direct competition with general intelligence labs.
- RAISE Summit18 min
Cheap Tokens, Expensive Mistakes: The Real Economics of AI at Scale | RAISE Summit 2026
Liran Zvibel, Dylan Patel, Karen Kwok
Major AI operators are pivoting from pure pricing power to structural cost efficiency, leveraging advanced KV cache offloading and optimized storage architectures to slash inference expenses. Recent benchmarks reveal that NAND-based memory solutions and strategic hardware combinations, such as high-capacity AMD GPUs, can outperform standard setups by up to 20 times in multi-turn agentic workflows while delivering 7x higher throughput. This shift allows emerging providers and frontier labs to significantly expand gross margins, with companies like Anthropic and OpenAI now projecting substantial operating profits driven by these infrastructure innovations rather than model weights alone.
- RAISE Summit18 min
Together AI & SemiAnalysis: In Conversation Together AI's Vision For the Future of AI Infrastructure
Vipul Vaid-Prakash, Dylan Patel
Founded in late 2022, Together AI challenges hyperscaler dominance by reengineering the cloud stack to achieve 75% maximum FLOP utilization on H100 GPUs, significantly outpacing the 50% industry standard. Led by chief scientist Tri Dao and the ClusterMax initiative, the firm drives down inference costs for massive models like DeepSeek from $8 to 30 cents per million tokens through architectural innovations and Reinforcement Learning frameworks. These technical advancements enable new MoE-based workloads and position the company to define emerging optimization surfaces before new hardware generations are fully adopted.