newsfilter.io

Savant Diaz

Showing 11 of 1 transcripts.

  1. Jane Street48 min

    Making GPUs Actually Fast: A Deep Dive into Training Performance

    Corwin, Savant Diaz, Sylvain De Wecker

    Jane Street engineers optimize deep learning infrastructure by eliminating CPU-GPU synchronization bottlenecks and fusing PyTorch operations via `torch.compile` and Triton to maximize throughput. When automated compilation fails on complex Python logic, they deploy custom CUDA kernels that leverage shared memory and warp-level reductions to achieve nearly 1,000x speedups in specialized tensor operations. This hierarchical approach, ranging from standard PyTorch to hand-optimized C++, ensures efficient utilization of the H100's 132 Streaming Multiprocessors and strict memory bandwidth constraints.