newsfilter.io
Other

What does the next training paradigm look like?

  • AGI is expected to be constructed through training AI across millions of verifiable tasks in thousands of diverse RL environments, with optimism that scaling will resolve current data inefficiency and lack of continual learning deficits.
  • Fundamental deficits in sample efficiency are anticipated to be amortized across billions of sessions, allowing AI agents to solve increasingly ambitious problems over longer time spans as RL training increases.
  • Architectural progress is predicted within a couple of years to yield context windows that effectively feel infinitely large, potentially rendering explicit continual learning of model weights unnecessary if in-context learning improves sufficiently.
  • Progress in computer use is expected to accelerate once AI agents achieve coding proficiency to build high-fidelity clones of applications like Slack and Gmail, provided replayable training targets for specific domains are created.
  • The generalization of RLVR is projected to create agents capable of planning, rapid learning, and skill acquisition within a single session, though short-horizon training may not generalize to long-horizon deployment.
  • While RLVR agents could theoretically provide superior strategic advice or build entities like SpaceX in historical contexts, a failure to transfer short-horizon learnings to long-horizon performance will prevent them from acquiring skills to build real-world businesses.
  • Without mechanisms to transfer session-based learnings into model weights, AI capabilities are predicted to be ephemeral, and current models will struggle to improve significantly in real-world interaction domains lacking replayable training targets.
  • Currently, 30% to 50% of lab compute is allocated to inference without contributing to model improvement, missing valuable deployment data regarding organizational context and failure modes.
  • Online learning is expected to remain viable only for limited use cases, as single session data is insufficient to train capable AIs for complex, specific jobs, and existing models likely continue learning the same objective across millions of users.
  • Sparse attention and KV cache compaction are not viewed as fundamental solutions to the continual learning bottleneck, with the loss function identified as the primary constraint that on-policy self-distillation (OPSD) may address.
  • A speculative "dreaming" approach is predicted to emerge as a fourth scaling axis, allowing AIs to practice endlessly against self-built simulations alongside pre-training, RL, and inference time compute.
  • By 2027 or 2028, agents are expected to be competent enough to gain real-world experience and distill learnings from a week of co-working with a human into their base models.
  • Future AI capabilities are anticipated to expand into adjacent domains via on-the-job learning once a continual learning recipe is established, shifting the primary mechanism of improvement from pre-release training to experience from broad economic deployment.
  • In the projected future state, every AI interaction will contribute to making models smarter not only through individual sessions but through the aggregation of interactions from all users globally.