newsfilter.io
Fireside Chat, Interview

New LLMs Are Unlocking Robot-Use Agents

Emerging Era of Robot Use Agents

  • Frontier researchers, including MIT's Philip Isola, suggest a shift toward "robot use agents" where general-purpose foundation models control diverse robotics embodiments.
  • Startups Waddle Labs and RoboCurve are leading the development of LLM-controlled robotics and physical AI evaluation, respectively.

Evolution from VLAs to Coding Agents

  • The RT2 paper established Vision-Language-Action (VLA) models that fine-tune pre-trained language models to output effector poses (robot joint coordinates) instead of text.
  • Early VLAs functioned like LLMs prior to the "Chain of Thought" revolution, executing actions directly without intermediate reasoning steps.
  • Modern coding agents utilize "code as policies," allowing models to write executable Python scripts rather than fixed action tokens, enabling complex tool creation and dynamic policy adaptation.
  • A key advantage of code-based approaches is "one-shot" learning; models can execute robot tasks immediately using their pre-existing coding knowledge without needing specific robotic fine-tuning data.

The Bitter Lesson and Data Modalities

  • The "Bitter Lesson" principle suggests that investing in general-purpose models trained on massive, diverse datasets yields better results than architecting specialized robotics models.
  • Models like Astra (a coding agent) leverage "computer use" data—such as interacting with GUIs, CAD software, and file systems—to learn spatial reasoning applicable to physical robotics.
  • This cross-modal transfer implies that training on digital environments effectively teaches models the topological and spatial concepts required for physical manipulation.
  • The industry is moving toward consolidating all available data types (coding, computer use, egocentric video) into single foundation models to maximize generalization.

Learning Paradigms: In-Context vs. Weight Updates

  • In-Context Learning (ICL) allows agents to adapt quickly to new tasks by appending examples to the context window, but performance caps out after roughly 20–40 examples due to context limits.
  • François Chollet distinguishes between "transduction" (slow, sample-heavy mapping) and "program induction" (generating a function from examples), noting that code acts as the efficient inductive bias for the latter.
  • Waddle Labs proposes a "harness" architecture that mimics biological memory consolidation: using the LLM for initial exploration and then compressing learned skills into faster, specialized programs for repeated use.
  • The framework suggests a hybrid approach where models "sleep" (offline) to distill in-context experiences into updated weights or compact skill libraries, similar to the DreamCoder algorithm.

Technical Demonstrations and Performance

  • Demo footage shows agents successfully performing multi-step tasks like unscrewing caps, uncapping pens, and manipulating blocks into bowls using camera inputs and tool calls.
  • Current end-to-end latency is a bottleneck, but latency for Frontier-class models is improving at a rate of approximately 2x per month.
  • Industry projections suggest real-time, general-purpose robot control will be economically viable by the end of the year.

Future Outlook and General Purpose Robotics

  • There is a consensus among frontier labs and robotics foundations that general-purpose robots capable of performing tasks like a competent teenager will emerge within two years.
  • The "Platonic Representation Hypothesis" posits that as models scale, their internal world representations converge, making strong language models inherently capable of robust physical control.
  • Future challenges involve managing growing libraries of skills and developing efficient data segregation frameworks to balance high-throughput execution with continuous learning.