newsfilter.io
Conference Presentation, Fireside Chat

Self-Improving Harnesses, Local Personal AI And YC's Agent For Work | YC Paper Club

  • Research on harnesses is shifting from being belittled to being recognized as worthy, with the static era ending in favor of self-improving harnesses developed over the last six months.
  • Leveraging test-time experience is expected to enable rapid adaptation to new domains, contrasting with current limitations where In-Context Learning saturates after 40 or 50 examples without parameter-efficient fine-tuning.
  • Future models are predicted to possess native reasoning loops and the ability to bootstrap themselves within a harness, potentially unlocking wild capabilities on the same weight file through self-improvement scaffolding.
  • Continual harness approaches, specifically dagger-style online learning capable of updating weight files, are anticipated to become a major research direction alongside the development of open systems like QM and OpenJarvis.
  • Evaluations such as Arc AGI and seven-day factorial runs involving 633 agents and 23 million tokens aim to demonstrate fluid intelligence, long-horizon performance, and the ability to maintain technology progression without plateauing.
  • Local LMs are expected to rival cloud counterparts on personal workflows and coding tasks within the near future, offering 800x lower costs as the gap closes through distillation and better accelerators.
  • The shift toward local inference on on-prem laptops and workstations is projected to account for a majority of daily inference calls, driven by cost efficiency and the need to avoid expensive H200 experiments.
  • Agentic systems are predicted to handle automated improvement loops and open problems using grind tools with walk clock time budgets, though current mixed results persist due to issues like "main character syndrome" and struggles with nuanced social context.
  • Trust in agent plans for database edits is expected to grow over the next few months, leading to employees increasingly rubber-stamping outputs, provided robust permission systems are in place for knowledge sharing.
  • Key takeaways for building harnesses include agentic context management, swarms, and standardized evals, with a focus on improving the cost-to-performance ratio by programmatically managing context.