newsfilter.io
Interview, Podcast

Open Models Are Collapsing The Cost Of AI

Market Trends and Adoption

  • Cost is the primary driver for enterprise adoption of open models, serving as an immediate solution that enables long-term goals of customization and control.
  • Open models now constitute the "North Star" for businesses seeking to customize AI for unique use cases rather than relying solely on proprietary closed models.
  • Interest in fine-tuning custom models, which waned earlier in 2024, is resurging as the release cadence of foundation models accelerates, making it harder to stay on top of generic updates.
  • AT&T has reportedly shifted 40% of its token consumption to open models, predominantly using US and European models while evaluating Chinese origins.
  • The surge in open model usage is driven by two distinct workflows: coding agents (developers) and AI assistants/co-work tools like OpenClaw and Hermes (non-developers in finance, support, and sales).
  • Per-developer token usage on Ollama Cloud saw an initial spike from coding agents (Kimi, GLM, Minimax) followed by an exponential growth in April due to the OpenClaw project enabling non-developers to automate complex tasks.
  • Ollama reported a 150x increase in aggregate token usage on its cloud platform since the start of the year, with per-user usage jumping from roughly 15 million to significantly higher volumes.
  • Open models have expanded context window capabilities from 128K to over 1 million tokens, enabling the explosive growth of long-span automated workflows.
  • Early 2024 saw a trend of serving off-the-shelf open models (e.g., DeepSeek, Kimi) as custom models, but late 2024/early 2025 marks a shift toward servicing out-of-the-box open models directly.

Geopolitics and Security

  • Security and safety remain the primary blockers for enterprise adoption of open models, particularly those of Chinese origin.
  • Despite geopolitical concerns, Fortune 500 companies indicate that if safety problems are solved, adoption of Chinese-origin models (hosted in secure environments like the US or Europe) is "completely on the table."
  • Open models are increasingly being used for cybersecurity tasks, such as pen-testing and vulnerability detection, where closed frontier models often refuse to assist.
  • There is a distinct difference between US and European customers (who prioritize model origin) and others (who prioritize where the model is hosted and run).
  • Open models are already powering critical infrastructure, such as Llama models used to detect power grid surges in Finland, highlighting the necessity of understanding data provenance.
  • The risk of "Manchurian Candidate" scenarios (booby-trapped models) is currently not verified, but companies rely on robust security teams to screen model weights and ensure supply chain integrity.

Ollama's Strategic Evolution and Technical Ecosystem

  • Ollama is used by 9 million developers, holds 178,000 GitHub stars, and serves 85% of the Fortune 500.
  • Ollama functions as an "operating system" for AI, orchestrating hardware drivers, inference engines, and application harnesses to provide a consistent developer experience.
  • The company's playbook for day-zero model launches involves validating hardware support, benchmarking against reference specs, and packaging models with compatible harnesses and documentation.
  • Ollama identified a pivot opportunity in 2023 when running Llama 2 locally revealed a severe lack of developer-friendly tools, leading to the creation of the Ollama platform.
  • Ollama raised a Series A in 2022 from Benchmark as a security-for-Kubernetes company before pivoting to local AI in 2023; they delayed monetization until the market matured for coding agents.
  • The company moved from a "walled garden" approach to building a "best-of-breed" ecosystem, addressing gaps in coordination, memory, and execution that frontier labs do not cover.
  • Ollama Cloud is currently dominated by Chinese-origin models for cloud-hosted use cases, whereas local deployments show a near-even split between US, European, and Chinese models.

Hardware and Deployment Models

  • A hybrid execution model is emerging where easier tasks (e.g., document processing) run locally on consumer hardware, while hard tasks (e.g., complex coding agents) route to cloud models.
  • Local hardware capabilities have advanced to run 20B to 120B parameter models; specifically, the 38B Qwen model now matches the coding performance of OpenAI's GPT-4.6.
  • NVIDIA's DGX Spark (with 128GB unified memory) and Apple Silicon are enabling desktop-level AI, allowing users to chain multiple units to run 400B+ models locally.
  • The "local-first" trend is expected to return to dominance for coding loops as hardware speeds approach the latency of cloud APIs (sub-100ms).
  • The scarcity of enterprise-grade GPUs (B200/B300) is driving the growth of inference providers like OpenRouter and Ollama to pool resources and simplify access for developers.

Economic Outlook and Future State

  • The projected steady state for enterprise AI involves 80-90% of token volume being processed by open models, though these may account for only 10-20% of the total budget.
  • Frontier closed models will remain reserved for the hardest tasks, while open models handle the "middle" and "grunt work" of workflows.
  • The future of AI agents relies on orchestration and chaining of smaller, ultra-low-cost "Flash" models (e.g., DeepSeek Flash) rather than reliance on a single "god-tier" model.
  • Open models are closing the performance gap to frontier closed models, now sitting only 2-3 months behind in intelligence and capability.
  • New business opportunities are emerging in "unbundled" layers such as knowledge injection (data context), coordination (agent routing), and execution (cloud sandboxes).
  • The industry is shifting from a "platform as a service" mindset to an "orchestration layer" mindset, where success depends on integrating fragmented models, hardware, and tools into a cohesive workflow.
  • Ollama is actively collaborating with hardware providers (NVIDIA, Apple) to optimize inference performance and ensure new models run efficiently on consumer devices.
  • The "unlimited token" experience, previously limited to closed providers like ChatGPT, is becoming viable for open models due to the extreme cost-efficiency of new architectures.