newsfilter.io

What If We Stopped Using GPUs? | YC Paper Club

Hardware-Software Co-Adaptation & The Backprop Bottleneck

  • Historical Context: From 2012 to 2020, AI hardware (specifically NVIDIA) optimized for "FLOPS per joule" to support Convolutional Neural Networks (CNNs), which were compute-bound rather than memory-bound.
  • The Transformer Shift (2020+): The emergence of GPT-2/3 shifted hardware priorities; Transformers are memory bandwidth and capacity bound due to the $O(N^2)$ complexity of attention mechanisms.
  • Efficiency Plateau: Gflops per joule improvements have stagnated in the last two years, necessitating a departure from current architectures.
  • Biological Baseline: The human brain achieves complex computation on ~20 watts using a massively feed-forward architecture, contrasting sharply with the energy-intensive backpropagation required by current silicon.
  • Backpropagation Impossibility: Biological neurons operate via feed-forward signals; backpropagation requires "weight transport" (identical weights for forward and backward paths) and backward signaling that effectively halts forward computation, neither of which occurs in the brain.
  • Hebbian Learning Challenges: Implementing Hebbian or similar rules in silico faces the "weight transport problem," requiring exact weight equality across distinct synaptic paths.
  • Inhibition Mechanisms: The brain utilizes cortical columns with massive inhibition to allow independent learning assemblies, a structural feature largely absent in standard transformer designs.

Alternative Optimizers: Zero-Order Methods

  • SPSA Method: The speaker's research focuses on SPSA (Simultaneous Perturbation Stochastic Approximation), a zero-order optimization method that estimates gradients via finite differences rather than backpropagation.
    • Mechanism: Perturbs model parameters in random directions ($\pm \epsilon$) and scales updates based on the resulting output difference.
    • Advantages: Capable of optimizing non-differentiable or discrete loss landscapes (e.g., Ackley's function) where gradient-based methods stall in local minima.
  • Scalability Limitations: Zero-order optimization becomes computationally infeasible for massive models because gradient noise scales linearly with model size; training a 10B parameter model would require an infeasible number of perturbations.
  • Sharded Optimization (SOMA): To mitigate noise, the "Sharded Optimization Mixture of Assemblies" (SOMA) strategy divides the model into small experts (e.g., 41K parameters) on separate GPUs.
    • Result: Sharding caps the gradient noise per expert, allowing convergence with only ~64 perturbations per step.
    • Scaling Laws: Performance improves as the model is sharded more aggressively, correlating with perturbations, batch size, and sharding depth.
  • Architecture Implications: Since gradients are not used, architectures designed for backpropagation (Transformers) are suboptimal; the speaker proposes LSTMs or custom feed-forward assemblies that function via pure forward passes.

Optical Computing

  • Physics Advantages: Photons exhibit ~10,000x lower loss and 10,000x higher bandwidth than electrons; they are inherently massively parallelizable as beams do not interact.
    • Energy Scaling: In optics, energy consumption scales linearly with input dimension ($N$), whereas in electronics, matrix-vector multiplication energy scales quadratically ($N^2$).
  • Key Bottlenecks:
    • I/O Conversion: Converting digital memory to optical signals (DAC) and back (ADC) consumes the vast majority of the system's energy budget (up to 99.99% of joules lost in conversion).
    • Non-Linearity: Passive optics only perform linear operations; implementing non-linear activation functions requires complex active mechanisms (e.g., short-pulse light in multimode fibers).
    • Storage: High-density, fast-write/read optical memory remains undeveloped compared to electronic NAND; optical storage is currently limited to archival (e.g., CD/DVD) or phase-change materials.
  • Application-Specific Approach: To overcome I/O costs, the speaker advocates for fixed-weight, application-specific optical accelerators where weights are microfabricated and never rewritten.
    • Diffusion Models: The speaker demonstrated an optical system programming light propagation to perform diffusion-based image generation (denoising) using a "digital twin" for calibration.
    • Weight Persistence: The system divides inference steps into sub-units where weights remain fixed for hundreds of steps to avoid reprogramming costs.
  • Manufacturing Constraints: Feature sizes are limited by the longer wavelength of light (micrometers vs. nanometers in silicon), though wavelength-division multiplexing can increase density.

Neuromorphic Computing

  • Core Principles: Neuromorphics seeks to emulate brain properties: intertwined memory/computation, event-driven "spiking" communication, and dynamic structural plasticity (rewiring via experience).
    • Spiking Neurons: Modeled as leaky capacitors that accumulate charge until a threshold is met, firing a discrete spike; a mix of analog (charge decay) and digital (discrete pulse) behavior.
  • Current State (2026 Outlook): The field remains in R&D, often struggling to balance biological inspiration with physical reality (hardware-software co-evolution).
  • Three Strategic Directions:
    1. Memory-Compute Co-location: Merging SRAM and logic (e.g., D-Matrix, IBM) to reduce data movement costs; often implemented in digital silicon but follows brain principles.
    2. Event-Driven/On-Device Learning: Using spiking networks for real-time edge sensing (e.g., Intel's drone chips) where the hardware adapts to local data.
    3. Exotic Substrates: Leveraging new physics (memristors, resistive memories, coupled oscillators) to co-design hardware and algorithms around specific nonlinear dynamics.
  • Biological vs. Silicon Limitations: While the brain is energy-efficient (20W), it lacks the clock speed and infinite storage capacity of digital computers; neuromorphic systems aim to hybridize these traits rather than perfectly mimic biology.
  • Noise Utilization: Emerging research (thermodynamic computing) explores using system noise for computation (e.g., Boltzmann machines), though ensuring reproducibility remains a significant challenge.

Biological Computation: "Human Cells in Doom"

  • Project Overview: Collaboration with Cortical Labs to train cultures of human brain cells (on micro-electrode arrays) to play the video game Doom.
  • Input/Output Mechanism:
    • Encoding: Game state (screens, ammo, health) is encoded into electrical stimuli (frequency, amplitude, channel selection) delivered to the cells.
    • Decoding: Neural spikes are read out via a linear decoder to map to game actions (move, aim, shoot).
  • Training Methodology:
    • RL Architecture: Uses Proximal Policy Optimization (PPO) with a "critic" network to guide learning.
    • Feedback Loop: Good actions trigger synchronous stimulation (low entropy); bad actions trigger asynchronous stimulation (high entropy).
    • Surprise Minimization: Feedback is scaled by the "surprise" (TD error) of the action relative to the critic's prediction, modulating the frequency and amplitude of the electrical feedback.
  • Avoiding "Cheat" Models: Critical to the success was preventing the silicon decoder from overfitting the task; the decoder was undersized to force the biological culture to generate the intelligence rather than the silicon simply playing the game.
  • Evolution of Intelligence: Unlike Pong (where hand-coded encoding/decoding sufficed), Doom required end-to-end learned encoding/decoding due to the high-dimensional action space.
  • Scalability & Ethics:
    • Scalability: Distributing biological intelligence to billions of users remains an unsolved hardware and networking challenge.
    • Optimization Goal: The cells optimize for reducing entropy (predictability) in the culture, not "feeling" pain; they are not human and do not possess consciousness or sentience.
    • Divergence: Biological compute substrates may evolve into a form of intelligence distinct from human cognition, optimized for the specific task rather than survival.