newsfilter.io
Interview, Fireside Chat

How Open Source Became AI's Backbone | Inferact with a16z

  • If GPU prices drop by 99%, the world may return to a real open source environment where individuals or small groups can train models in basements or during free time.
  • If moderation remains unsolved, users are predicted to default to open models by 2030 to control guardrails for trusted use cases, as open models may finally close the capability gap with frontier models by then.
  • Hardware vendors including NVIDIA, AMD, Google, Amazon, and Intel plan to ensure new chips run VLLM software via a multi-party co-design process, often using VLLM as a benchmark to bridge a 10x performance gap.
  • Proprietary model providers plan to offer only regular and fast modes, whereas open weight model providers can potentially offer up to 10 different speed levels.
  • Users running open weight models may access up to 10 speed tiers ranging from slowest modes to 400-500 tokens per second, which is typically 2x or 3x faster than the fast mode available for proprietary models.
  • The VLLM team plans to enable fast modes allowing developers to execute environment interactions asynchronously rather than being stuck in sync.
  • Model labs are expected to increasingly add economic terms to licenses, similar to Meta's Llama threshold approach, to fund massive capital expenditures for training and R&D.
  • Open source models are anticipated to serve as a necessary economic structure for funding, likened to the pharmaceutical industry's R&D funding model, due to the enormous compute required to reach the AI frontier.
  • Community effort is expected to continue adapting open models to diverse cluster topologies, edge devices, and specialized use cases such as voice and coding agents.
  • Open source inference is expected to remain the leading method for running models, with many inference clouds and API services leveraging open source engines under the hood.
  • Users expect the ability to fine-tune open models to improve performance for specific workloads, a practice already observed with models like Kimi K3.
  • Future innovation is predicted to depend on the training environment rather than just data, making the distillation of the learning process difficult or impossible to replicate.
  • The "village" of community effort is expected to validate and fix rare bugs occurring 0.0001% of the time while optimizing models against an almost infinite variety of use cases.
  • Global progress relies on the "magic" of collaboration occurring when people engage in more open source model training globally rather than in specific regions.
  • Discourse on new frontier models is expected to focus on the ability to own, run, fine-tune, and understand the performance profile rather than solely on cost comparisons.
  • Cost will continue to be a factor for users regarding migration off expensive plans, though control over infrastructure and system performance remains a backbone requirement.
  • Open source models allow for the calibration of speed and performance, giving users control over data retention and security compliance.