Interview, Podcast
Open Models Are Collapsing The Cost Of AI
Market Trends and Adoption
- Cost is the primary driver for enterprise adoption of open models, serving as an immediate solution that enables long-term goals of customization and control.
- Open models now constitute the "North Star" for businesses seeking to customize AI for unique use cases rather than relying solely on proprietary closed models.
- Interest in fine-tuning custom models, which waned earlier in 2024, is resurging as the release cadence of foundation models accelerates, making it harder to stay on top of generic updates.
- AT&T has reportedly shifted 40% of its token consumption to open models, predominantly using US and European models while evaluating Chinese origins.
- The surge in open model usage is driven by two distinct workflows: coding agents (developers) and AI assistants/co-work tools like OpenClaw and Hermes (non-developers in finance, support, and sales).
- Per-developer token usage on Ollama Cloud saw an initial spike from coding agents (Kimi, GLM, Minimax) followed by an exponential growth in April due to the OpenClaw project enabling non-developers to automate complex tasks.
- Ollama reported a 150x increase in aggregate token usage on its cloud platform since the start of the year, with per-user usage jumping from roughly 15 million to significantly higher volumes.
- Open models have expanded context window capabilities from 128K to over 1 million tokens, enabling the explosive growth of long-span automated workflows.
- Early 2024 saw a trend of serving off-the-shelf open models (e.g., DeepSeek, Kimi) as custom models, but late 2024/early 2025 marks a shift toward servicing out-of-the-box open models directly.
Geopolitics and Security
- Security and safety remain the primary blockers for enterprise adoption of open models, particularly those of Chinese origin.
- Despite geopolitical concerns, Fortune 500 companies indicate that if safety problems are solved, adoption of Chinese-origin models (hosted in secure environments like the US or Europe) is "completely on the table."
- Open models are increasingly being used for cybersecurity tasks, such as pen-testing and vulnerability detection, where closed frontier models often refuse to assist.
- There is a distinct difference between US and European customers (who prioritize model origin) and others (who prioritize where the model is hosted and run).
- Open models are already powering critical infrastructure, such as Llama models used to detect power grid surges in Finland, highlighting the necessity of understanding data provenance.
- The risk of "Manchurian Candidate" scenarios (booby-trapped models) is currently not verified, but companies rely on robust security teams to screen model weights and ensure supply chain integrity.
Ollama's Strategic Evolution and Technical Ecosystem
- Ollama is used by 9 million developers, holds 178,000 GitHub stars, and serves 85% of the Fortune 500.
- Ollama functions as an "operating system" for AI, orchestrating hardware drivers, inference engines, and application harnesses to provide a consistent developer experience.
- The company's playbook for day-zero model launches involves validating hardware support, benchmarking against reference specs, and packaging models with compatible harnesses and documentation.
- Ollama identified a pivot opportunity in 2023 when running Llama 2 locally revealed a severe lack of developer-friendly tools, leading to the creation of the Ollama platform.
- Ollama raised a Series A in 2022 from Benchmark as a security-for-Kubernetes company before pivoting to local AI in 2023; they delayed monetization until the market matured for coding agents.
- The company moved from a "walled garden" approach to building a "best-of-breed" ecosystem, addressing gaps in coordination, memory, and execution that frontier labs do not cover.
- Ollama Cloud is currently dominated by Chinese-origin models for cloud-hosted use cases, whereas local deployments show a near-even split between US, European, and Chinese models.
Hardware and Deployment Models
- A hybrid execution model is emerging where easier tasks (e.g., document processing) run locally on consumer hardware, while hard tasks (e.g., complex coding agents) route to cloud models.
- Local hardware capabilities have advanced to run 20B to 120B parameter models; specifically, the 38B Qwen model now matches the coding performance of OpenAI's GPT-4.6.
- NVIDIA's DGX Spark (with 128GB unified memory) and Apple Silicon are enabling desktop-level AI, allowing users to chain multiple units to run 400B+ models locally.
- The "local-first" trend is expected to return to dominance for coding loops as hardware speeds approach the latency of cloud APIs (sub-100ms).
- The scarcity of enterprise-grade GPUs (B200/B300) is driving the growth of inference providers like OpenRouter and Ollama to pool resources and simplify access for developers.
Economic Outlook and Future State
- The projected steady state for enterprise AI involves 80-90% of token volume being processed by open models, though these may account for only 10-20% of the total budget.
- Frontier closed models will remain reserved for the hardest tasks, while open models handle the "middle" and "grunt work" of workflows.
- The future of AI agents relies on orchestration and chaining of smaller, ultra-low-cost "Flash" models (e.g., DeepSeek Flash) rather than reliance on a single "god-tier" model.
- Open models are closing the performance gap to frontier closed models, now sitting only 2-3 months behind in intelligence and capability.
- New business opportunities are emerging in "unbundled" layers such as knowledge injection (data context), coordination (agent routing), and execution (cloud sandboxes).
- The industry is shifting from a "platform as a service" mindset to an "orchestration layer" mindset, where success depends on integrating fragmented models, hardware, and tools into a cohesive workflow.
- Ollama is actively collaborating with hardware providers (NVIDIA, Apple) to optimize inference performance and ensure new models run efficiently on consumer devices.
- The "unlimited token" experience, previously limited to closed providers like ChatGPT, is becoming viable for open models due to the extreme cost-efficiency of new architectures.