Interview, Podcast
Context Engineering Our Way to Long-Horizon Agents: LangChain’s Harrison Chase
- Traces are projected to exert greater impact on agent systems than single LLM applications due to uncertainty in later steps caused by preceding loops with arbitrary inputs.
- Long-horizon agents are expected to lead in the coding domain first, with capabilities expanding to other sectors as models improve and scaffolding enhances.
- Agents are anticipated to operate for increasingly long durations producing "first drafts" in finance research and customer support that still require human review.
- The "Opus 4.5" release around November and December marked a period where coding agents became significantly more capable of handling long-horizon tasks.
- The industry is expected to transition from scaffolds to harnesses in early 2025 or late 2025 once models reach a threshold where they can orchestrate their own loops.
- Future iterations will likely grant LLMs "context engineering" authority to compact data, expanding on current experimental features.
- Models are expected to continuously improve at longer-horizon tasks, potentially reducing the need for manual scaffolding over time.
- A general-purpose agent is likely to manifest as a "coding engine," whereas current coding agents remain optimized for specific tasks.
- Building agents will require significantly more iteration than traditional software development because exact behavior cannot be predicted prior to shipping.
- Existing software companies with valuable data are expected to easily integrate agent systems to generate real value if necessary instructions are defined.
- Most users are predicted not to build their own harnesses in the long run due to complexity, leading to reliance on external providers.
- Human-in-the-loop evaluation will remain critical, with "LLM as a judge" capabilities evolving to calibrate grading traces against human preferences.
- Coding agents are expected to increasingly utilize tools like "Langsmith MCP" and "Langsmith Fetch" for automated trace diagnosis and code correction.
- Future development may include "dreaming" or "sleep time compute," allowing agents to run nightly tasks to review traces and update instructions autonomously.
- Memory capabilities are expected to serve as a significant moat for specific workflow agents, though they are not predicted to increase stickiness for general chat models like ChatGPT.
- User interfaces for managing long-horizon agents are expected to evolve to include both "sync" and "async" modes for reviewing outputs and providing feedback.
- Future interfaces are projected to feature an "agent inbox" enabling users to enter sync mode to chat and view modified states, such as file system changes.
- Access to code execution and file systems is considered essential for the "longer tail of use cases," with "90%" agreement on the necessity of coding.
- Browser use by agents is not expected to be viable currently due to model limitations, though it may eventually be approximated via coding agents using a CLI.
- Predictions regarding the field are acknowledged as difficult, with a recognition that future assessments may contradict current statements.