newsfilter.io
Conference Presentation, Keynote

A Keynote by Thomas Sohmers, Co Founder & CTO of Positron

  • The AI inference explosion is projected to become the primary bottleneck for global economic growth, fundamentally reshaping the macroeconomic model and justifying massive capital and operational expenditures on data center expansion.
  • A reduction in effective inference costs from approximately $7 per hour to under $1 is expected to trigger widespread trickle-down economic effects, while the scaling of labor capabilities creates a dual impact of new growth opportunities and job displacement.
  • Future training processes will increasingly be bottlenecked by inference requirements due to reasoning methods like Monte Carlo Tree Search, necessitating the development of specialized training systems with distinct underlying needs.
  • Positron's hardware strategy focuses on optimizing performance per dollar and per watt, with current products delivering a 70% speed advantage over two NVIDIA H100 devices, less than half the retail cost, and significantly lower power consumption.
  • Deployment capabilities are designed to fit air-cooled data centers that cannot accommodate the massive cooling requirements of competitor solutions, utilizing 150-watt PCI cards and ensuring drop-in compatibility with the NVIDIA CUDA ecosystem via direct ingestion of Hugging Face Transformers files.
  • The upcoming "Atlas" server is expected to handle 512 billion parameter models concurrently to resolve low utilization rates for service providers hosting multiple models, while the second-generation "Titan" system is committed to shipping next year with next-generation silicon.
  • The Titan system is projected to support model sizes reaching many trillions of parameters by attaching one terabyte of DRAM to each chip, a configuration described as impossible on current GPU or alternative platforms.
  • The long-term outlook posits a technological shift where the future of AI becomes "positronic rather than electronic," driven by these advancements in inference economics and hardware efficiency.