newsfilter.io
Fireside Chat, Interview

New LLMs Are Unlocking Robot-Use Agents

  • Frontier researchers anticipate entering an era where general-purpose models enable different robots to become capable through Vision-Language-Action models and coding agents that control hardware out of the box without specific fine-tuning.
  • Current progress in these models is constrained by data availability rather than architecture, with experts identifying the transfer of robotics data into language model distributions as a potential pathway to unlock further domain capabilities.
  • Future models are expected to significantly improve at tool use, allowing them to write complex policies as code and explore environments via in-context learning, potentially shifting the industry toward single foundational Large Language Models instead of robot-specific training.
  • Investing compute into coding capabilities is predicted to trigger a "liftoff" where AI agents automate machine learning engineering and displace roles in robotics, coding, and art.
  • In-context learning is projected to saturate after approximately 20 to 40 examples due to context window limits, necessitating the use of Retrieval Augmented Generation or active memory for scaling beyond these thresholds.
  • A hierarchy of learning paradigms ranging from in-context learning to Supervised Fine-Tuning and Reinforcement Learning exists, with in-context learning proving most effective in low-data regimes while uncertainty remains regarding the boundary between weight space and symbolic space.
  • Future systems are expected to consolidate in-context skills into specific programs or tools within a harness to distill past experience, alongside a "sleep" phase for data segregation, weight updates, and skill refactoring to manage growing libraries of deployment data.
  • Latency for frontier-class Large Language Models is improving by approximately 2x per month, with real-time control predicted to be achievable by the end of the year to overcome economic utility barriers.
  • Coding agents are expected to utilize loops with Vision-Language Models to handle variation and failure cases while maintaining a deterministic core policy graph.
  • The convergence of coding, computer-use, and egocentric video data into a single model is expected to yield the most capable robot use agents, driven by the "Platonic Representation Hypothesis" where strong language and robotics models converge on similar world representations.
  • Pre-training on extensive computer-use data, including CAD files and GUI interactions, is expected to teach spatial reasoning that transfers effectively to physical robot control.
  • Designing robot interfaces that mimic computer-use tools, such as drag-and-drop cursors, is expected to improve LLM performance on physical tasks.
  • Consensus among frontier labs and robotics foundation model companies indicates that general-purpose robots will exist within the next two years or earlier, capable of performing any task a competent teenager could do with their bare hands upon receiving a natural language instruction.
  • The Waddle Labs team expects to take steps toward general-purpose robot capabilities in about two years' time.
  • Society is expected to be largely unaware or unprepared for the capabilities of general-purpose robots arriving within the next two years, though awareness of the extent of upcoming advancements is anticipated over the following few years.