newsfilter.io
Interview, Podcast

Context Engineering Our Way to Long-Horizon Agents: LangChain’s Harrison Chase

  • Traces are projected to exert greater impact on agent systems than single LLM applications due to uncertainty in later steps caused by preceding loops with arbitrary inputs.
  • Long-horizon agents are expected to lead in the coding domain first, with capabilities expanding to other sectors as models improve and scaffolding enhances.
  • Agents are anticipated to operate for increasingly long durations producing "first drafts" in finance research and customer support that still require human review.
  • The "Opus 4.5" release around November and December marked a period where coding agents became significantly more capable of handling long-horizon tasks.
  • The industry is expected to transition from scaffolds to harnesses in early 2025 or late 2025 once models reach a threshold where they can orchestrate their own loops.
  • Future iterations will likely grant LLMs "context engineering" authority to compact data, expanding on current experimental features.
  • Models are expected to continuously improve at longer-horizon tasks, potentially reducing the need for manual scaffolding over time.
  • A general-purpose agent is likely to manifest as a "coding engine," whereas current coding agents remain optimized for specific tasks.
  • Building agents will require significantly more iteration than traditional software development because exact behavior cannot be predicted prior to shipping.
  • Existing software companies with valuable data are expected to easily integrate agent systems to generate real value if necessary instructions are defined.
  • Most users are predicted not to build their own harnesses in the long run due to complexity, leading to reliance on external providers.
  • Human-in-the-loop evaluation will remain critical, with "LLM as a judge" capabilities evolving to calibrate grading traces against human preferences.
  • Coding agents are expected to increasingly utilize tools like "Langsmith MCP" and "Langsmith Fetch" for automated trace diagnosis and code correction.
  • Future development may include "dreaming" or "sleep time compute," allowing agents to run nightly tasks to review traces and update instructions autonomously.
  • Memory capabilities are expected to serve as a significant moat for specific workflow agents, though they are not predicted to increase stickiness for general chat models like ChatGPT.
  • User interfaces for managing long-horizon agents are expected to evolve to include both "sync" and "async" modes for reviewing outputs and providing feedback.
  • Future interfaces are projected to feature an "agent inbox" enabling users to enter sync mode to chat and view modified states, such as file system changes.
  • Access to code execution and file systems is considered essential for the "longer tail of use cases," with "90%" agreement on the necessity of coding.
  • Browser use by agents is not expected to be viable currently due to model limitations, though it may eventually be approximated via coding agents using a CLI.
  • Predictions regarding the field are acknowledged as difficult, with a recognition that future assessments may contradict current statements.