Fireside Chat, Interview
New LLMs Are Unlocking Robot-Use Agents
Emerging Era of Robot Use Agents
- Frontier researchers, including MIT's Philip Isola, suggest a shift toward "robot use agents" where general-purpose foundation models control diverse robotics embodiments.
- Startups Waddle Labs and RoboCurve are leading the development of LLM-controlled robotics and physical AI evaluation, respectively.
Evolution from VLAs to Coding Agents
- The RT2 paper established Vision-Language-Action (VLA) models that fine-tune pre-trained language models to output effector poses (robot joint coordinates) instead of text.
- Early VLAs functioned like LLMs prior to the "Chain of Thought" revolution, executing actions directly without intermediate reasoning steps.
- Modern coding agents utilize "code as policies," allowing models to write executable Python scripts rather than fixed action tokens, enabling complex tool creation and dynamic policy adaptation.
- A key advantage of code-based approaches is "one-shot" learning; models can execute robot tasks immediately using their pre-existing coding knowledge without needing specific robotic fine-tuning data.
The Bitter Lesson and Data Modalities
- The "Bitter Lesson" principle suggests that investing in general-purpose models trained on massive, diverse datasets yields better results than architecting specialized robotics models.
- Models like Astra (a coding agent) leverage "computer use" data—such as interacting with GUIs, CAD software, and file systems—to learn spatial reasoning applicable to physical robotics.
- This cross-modal transfer implies that training on digital environments effectively teaches models the topological and spatial concepts required for physical manipulation.
- The industry is moving toward consolidating all available data types (coding, computer use, egocentric video) into single foundation models to maximize generalization.
Learning Paradigms: In-Context vs. Weight Updates
- In-Context Learning (ICL) allows agents to adapt quickly to new tasks by appending examples to the context window, but performance caps out after roughly 20–40 examples due to context limits.
- François Chollet distinguishes between "transduction" (slow, sample-heavy mapping) and "program induction" (generating a function from examples), noting that code acts as the efficient inductive bias for the latter.
- Waddle Labs proposes a "harness" architecture that mimics biological memory consolidation: using the LLM for initial exploration and then compressing learned skills into faster, specialized programs for repeated use.
- The framework suggests a hybrid approach where models "sleep" (offline) to distill in-context experiences into updated weights or compact skill libraries, similar to the DreamCoder algorithm.
Technical Demonstrations and Performance
- Demo footage shows agents successfully performing multi-step tasks like unscrewing caps, uncapping pens, and manipulating blocks into bowls using camera inputs and tool calls.
- Current end-to-end latency is a bottleneck, but latency for Frontier-class models is improving at a rate of approximately 2x per month.
- Industry projections suggest real-time, general-purpose robot control will be economically viable by the end of the year.
Future Outlook and General Purpose Robotics
- There is a consensus among frontier labs and robotics foundations that general-purpose robots capable of performing tasks like a competent teenager will emerge within two years.
- The "Platonic Representation Hypothesis" posits that as models scale, their internal world representations converge, making strong language models inherently capable of robust physical control.
- Future challenges involve managing growing libraries of skills and developing efficient data segregation frameworks to balance high-throughput execution with continuous learning.