Interview, Fireside Chat
Building the Automated AGI Lab: Core Automation's Jerry Tworek and Rohan Anil
- Core Automation plans to replace the transformer architecture within a "year or two" by developing alternatives not currently profitable or immediate winners, aiming to redefine the field and achieve meaningful long-term adaptability at test time.
- The company intends to accelerate the transition by automating the full deep learning stack, including kernel generation, to reduce the time from idea to execution and potentially reach iteration speeds of 10 to 200 operations per day for search and optimization.
- Researchers anticipate that biological learning cannot be surpassed with current hardware and will require new hardware designs, such as analog circuits, to match human operation efficiency.
- The team predicts that the market will likely resist adopting non-profitable alternative architectures for a "much longer time horizon" due to the immediate profitability of transformer scaling, necessitating Core Automation's intervention to force change.
- Current transformers are expected to be "capped out" in their ability to handle new relationships, tasks, or code bases without retraining, becoming "less and less useful" over time if not continuously fed new world events and data.
- Achieving AGI requires models to improve themselves without human loops, necessitating a shift from current methods which rely heavily on human-in-the-loop optimization for high-performance tasks like 60x faster QR operations.
- The organization aims to be the "most automated lab there is," granting researchers maximum agency while testing system autonomy by evaluating performance during extended team vacations to determine if the lab can produce better results without human intervention.
- Pre-training and reinforcement learning are viewed as distinct optimization problems that, when combined, could yield an order of magnitude improvement, though current scaling laws have made alternative approaches like LSTMs economically unviable.
- Fundamental research in deep learning historically takes five to six years to reach industry adoption, with recent developments like sparsity and mixtures of experts requiring two to three years to refine post-invention.
- Kernels remain a critical bottleneck for training and optimization, with few experts possessing the skills to write high-performance code, leading to a significant efficiency gap between current automated solutions and human-optimized implementations.
- The team believes that inference time scaling and speculative decoding are insufficient band-aids, and that true efficiency gains will come from end-to-end thinking that integrates optimization methods with architectural design.
- Evaluating system success involves looking for performance "plots" where multiple "all right pieces fall into it," with a specific benchmark being the ability to solve complex matrix factorization problems without human assistance.
- The current path is deemed insufficient for AGI, prompting a focus on "serious research" to unlock test-time learning and meta-learning on the architectural layer for horizons longer than what is currently achievable.
- Many architectures require a baseline of compute to function effectively, suggesting that small-scale research may be insufficient to discover superior systems capable of test-time adaptation.
- The company expects that the majority of computational efficiency will be found in systems that can learn continuously from deployment, contrasting with the "notably difficult" nature of removing humans from the learning loop.
- Transformers are currently limited by low data efficiency and catastrophic forgetting during fine-tuning, whereas biological learning involves multiple algorithms working together rather than single-process reinforcement learning.
- The team anticipates that as hardware and algorithms evolve, the industry will eventually accept spending less compute for better efficiency if the performance gains are realized through new architectural discoveries.
- Deep learning systems require five specific components to align correctly before a successful performance curve is observed, and finding the right architecture often involves a "click" moment where previously unknown pieces integrate.
- Human-LLM hybrids are currently highly successful, but fully autonomous models remain "not at all" successful, highlighting the critical challenge of removing human intervention from the operational loop.