newsfilter.io
Interview, Fireside Chat

Open Models vs Frontier Models: Who Actually Wins? | The $100K Token Budget Every Engineer Will Need

  • Unbounded demand for frontier intelligence is expected to persist as open models face limits in high-stakes domains, while open weights models will eventually handle most automated enterprise tasks, though this represents a minor fraction of total volume.
  • Frontier models like GPT-4 are projected to become commoditized at 1/300th the current cost for equivalent intelligence, creating an assembly line approach with fine-tuned models, yet hardware improvements may be offset by energy constraints establishing a cost floor for token usage.
  • Token budgeting will evolve into a standard practice where CFOs allocate capital per employee, with token spend converging to 20% of developer salaries from the current 3.8%, eventually reaching $100,000 annually per top engineer.
  • The AI market is predicted to shift from a cloud infrastructure model to an "Uber-Lyft" dynamic driven by scale and industry depth, with Sierra aiming to be a leading competitor in this space.
  • Coding agents are forecasted to drive a fundamental step change in software development within six weeks to months, shifting the bottleneck from code writing to code review and deciding what is worth building.
  • Customer deployment timelines are accelerating due to forward-deployed teams, with cases like "Next" going live in six weeks and "Cigna" in 58 days, while cybersecurity is expected to see increased offensive capabilities and demand.
  • Young graduates are anticipated to hold an unfair advantage over experienced workers due to early mastery of AI tools, necessitating a focus on rapid learning and intensity.
  • Sierra plans to expand its forward-deployed team to Japan by acquiring Opera Technologies to hit the ground running this year, eventually deploying 10 people on the ground to meet local "omotenashi" standards.
  • The company intends to build domain expertise in verticals like retail and healthcare, scaling lessons from unique Fortune 50 to Fortune 10 deployments to reduce the prevalence of "one of one" solutions.
  • Sierra expects 50% of its customers to generate over $1 billion in revenue and 30% to exceed $10 billion, targeting a future where token spend is budgeted alongside headcount as the company moves from learning-focused discipline to capital efficiency.
  • Expansion goals include reaching 100 employees in Europe soon and developing teams attuned to cultural nuances to support global operations.
  • Sierra plans to evolve from customer support into complete lifecycle management, becoming an inbound sales machine, while continuing to build custom solutions for unique customer needs like cart abandonment features.
  • "Sierra Brain" is projected to become an indispensable strategy thought partner and operational tool, incorporating company data, board letters, and operating reviews to reason about company direction, support hiring reviews via "Clay scanner," and provide employees with an "MCP gateway" to interrogate the entire organization.
  • Future iterations of "Sierra Brain" will include a 20-30 page grounding document on company structure, access to all documents, Slack messages, presentations, and operating reviews, and will be used to build private agents and speed up software development.
  • The organization will maintain a six-week board cadence to adapt to the fast-moving AI landscape rather than a quarterly schedule, while continuing to be selective about engagement levels to preserve intensity and avoid hands-off management.
  • Sierra will continue to build applications to inform its platform, develop agents including "Ghostwriter," and avoid relying solely on external hiring, ensuring a "furnace around the core" of the platform.
  • Cybersecurity is viewed as a strong growth area due to ratcheting up offensive capabilities, though uncertainty remains regarding whether specific tools like "Mythos" or "Codex 55 Cyber" will become the primary solutions.
  • Local model execution on phones is expected to improve consumer applications but will not alleviate the server-side compute requirements for training and complex inference involving petaflops or exaflops.