newsfilter.io
Interview, Podcast

Open Models Are Collapsing The Cost Of AI

  • Cost is anticipated to remain a short-term constraint that ultimately enables businesses to customize models for unique use cases, while early 2024 fine-tuning efforts are predicted to be rendered obsolete by subsequent model releases.
  • Open source and open weight models are expected to continue improving and eventually bridge the intelligence gap with closed frontier models within less than three months, causing the frontier labs to slow their release cadence to address alignment and containment.
  • The market is projected to reach a steady state where 80 to 90 percent of tokens are served by open models, which will be funded by only 10 to 20 percent of the budget due to lowered costs, while the most complex tasks remain reserved for frontier labs.
  • A new class of "Flash models," led by initiatives like DeepSeek, is expected to become the workhorse for "grunt work" and high-volume token usage, enabling unlimited tokens and making users indifferent to underlying per-token costs when chaining models.
  • Local hardware capabilities are expected to expand rapidly, with next-generation devices like GB300 and DGX stations enabling consumer and desktop environments to run 20 to 120 billion parameter models, returning the coding loop to local execution with latency comparable to running tests.
  • The industry is expected to shift toward an orchestration-heavy model where a "router" or frontier model manages scheduling and harder tasks, while simpler tasks are handled locally, creating a blend of US, European, and Chinese origin models for different deployment contexts.
  • Geopolitical factors and security concerns regarding model origin are expected to drive a split in the market, with some customers prioritizing the source of the model for mission-critical tasks while others focus on where the model is hosted and whether supply chain poisoning is mitigated through safety checks.
  • Supply chain volatility is expected to impact the availability of high-end GPUs like B200 and B300, making it difficult for startups to access these resources, while the scarcity of model tokens will be replaced by a new scarcity in orchestration and efficient model integration.
  • The next generation of hardware and software will support a renaissance for personal desktops, featuring specialized accelerators for image and audio processing, and enabling the running of large models via device chaining or miniature data center racks.
  • Monetization strategies are expected to pivot toward coding agents initially, expanding to non-developers through projects like OpenClaw and Hermes, with the cloud growth driven by the transition from hobbyist adoption to enterprise integration where permissionless access is a key driver.
  • The "god model" concept is not expected to solve the majority of use cases, as the industry will move toward a mix of smaller, specialized, and efficient models that handle the vast majority of tasks, reserving larger models for unlocking specific difficult problems.
  • Software maintenance and security are expected to remain critical challenges, with the "rules" of previous infrastructure eras breaking down in the AI world, requiring robust safety training and continuous updates to ensure long-term functionality and prevent supply chain poisoning.
  • The gap between US and Chinese model performance is expected to close significantly, with Chinese models currently driving cloud usage growth and expected to be predominantly consumed in the cloud, while a strong blend of origins will occur locally.
  • The Ollama platform is expected to see its cloud growth driven by coding agents and the ability to solve workflow problems, leveraging a "first-to-market" developer experience to facilitate the adoption of open models in both local and cloud environments.
  • Future startups are expected to be formed by teams that have achieved product-market fit, utilizing the "muscle memory" of previous successful infrastructure companies to navigate the unique challenges of AI software, which requires holding customers accountable to quality standards that AI alone cannot enforce.