Interview, Fireside Chat, Conference Presentation
Jane Street on GPUs, Trading, and Hiring: A Conversation with Dwarkesh
- Jane Street expects to leverage a $6 billion compute agreement with CoreWeave to train diverse model architectures, with plans to scale GPU capacity from tens of thousands to hundreds of thousands within an unspecified short timeframe to reduce iteration times and drive innovation.
- Infrastructure strategy involves moving away from single-architecture CPU reliance and monolithic data centers toward diverse, modular setups that support both ARM and x86-64 architectures, utilizing off-site built components shipped for plug-and-play deployment to navigate over one-year lead times for items like generators.
- Physical constraints such as power density requiring one-megawatt racks, liquid cooling requirements for temperatures lower than current NVIDIA H100 standards, and the difficulty of wiring sufficient power into single locations will necessitate spreading compute across multiple facilities rather than using single large clusters.
- The organization anticipates that human talent acquisition, specifically the time required to train new hires and establish mentorship capacity, will be the primary bottleneck for growth rather than hardware availability, driving continued hiring in mechanical/electrical engineering, project management, and machine learning.
- Investment justification for massive compute expenditures is based on long-term research value and business potential rather than immediate P&L metrics or specific trading strategies, with the "dark space" of unrealized value due to compute constraints viewed as a significant opportunity.
- System design priorities will shift toward handling high-volume, noisy financial data with a different bytes-to-flops ratio, requiring robust internal object stores, emphasis on data loading performance, and the use of FPGAs for sub-100-nanosecond trading regimes where hardware constraints override programming language choices.
- Jane Street plans to build a formal methods team using mathematical proofs to improve software engineering effectiveness and will utilize reserve compute capacity for model retraining to counteract quality decay and bulk inference tasks to fill scheduling gaps.
- Future competitiveness relies on investing in difficult-to-automate trading areas, such as bond markets and human judgment during "phase transitions" or market volatility, while maintaining that current AI systems have not reached an AGI level that would strictly outperform human meta-judgment in complex scenarios.
- The firm expects the rapid obsolescence of data center components, with infrastructure details potentially becoming outdated within two weeks, necessitating agile procurement practices like stocking fungible components and making infrastructure decisions before chip orders due to the critical nature of generator lead times.
- Competitive risks include the potential for rivals to develop methods that reduce the value of current work, the possibility that model and trade value appreciation may lag expectations, and the physical bottlenecks of securing colocation space and fiber latency optimization for ultra-fast trading.