newsfilter.io
Interview, Fireside Chat, Conference Presentation

How to Build Your Own Data Center & Why Every Startup Should Do It

Strategic Pivots and Market Positioning

  • Cliff Weitzman admits 11 Labs "leapfrogged" Speechify in the B2B space, labeling his decision to avoid the B2B API model initially as "the biggest strategic mistake" in the company's history.
  • Weitzman previously believed text-to-speech APIs would become commoditized and run locally on devices, failing to anticipate the "compound startup" model where an initial wedge product (API) evolves into a broader ecosystem (AI agents, voice cloning, duplex models).
  • Speechify now competes directly with 11 Labs and Sierra, having launched products to rival Siri and 11 Labs' offerings after recognizing that "you cannot not be in the race."
  • Speechify dominates the B2C market with 98% of app store installs for text-to-speech and has processed over 770 billion words (approx. 6,000 years of listening time) to date.
  • Weitzman distinguishes between "generalist" and "specialist" AI labs, asserting that specialized models trained on proprietary data will become the norm for enterprises.
  • Weitzman notes that 11 Labs has achieved significant government buy-in across major Western democracies within three months of launching, whereas Speechify focuses on a different growth trajectory.

Hardware Strategy and Compute Economics

  • Speechify invests tens of millions of dollars in owning NVIDIA GPUs (including early access to H100s, B200s, and Rubens) rather than renting exclusively, citing a 1.5x cost efficiency where owning hardware is cheaper than renting for one year.
  • Owning hardware allows engineers to work without "parsimony," removing the fear of costing the company thousands of dollars per hour during training iterations.
  • Weitzman argues that co-located memory in large GPU clusters is essential for large-scale training, a requirement that makes renting from hyperscalers (AWS, Azure, GCP) inefficient for Speechify's specific workload.
  • The company utilizes a hybrid asset strategy: purchasing hardware for baseline utilization (training and high-volume inference) while renting spot instances and dedicated instances from hyperscalers to handle demand spikes (e.g., September back-to-school surge).
  • Legacy GPUs (e.g., K80s, A100s) are retained for inference tasks, decoupling "training" speed requirements from "inference" quality requirements, extending the useful life of older hardware.
  • Speechify pays a $100,000 premium per GPU to receive shipments four months early to bypass supply chain queues and secure "Rubens" liquid-cooled units before competitors.
  • Weitzman highlights NVIDIA's new underwriting deals with Blackstone, BlackRock, Apollo, and Goldman Sachs, which guarantee 25% of GPU value, creating a liquid secondary market and reducing depreciation risk.
  • Data center constraints are identified as physical space, energy (the primary constraint), and liquid cooling; Speechify utilizes side-car liquid cooling units to meet the thermal demands of Blackwell architectures.

Product Development and AI Engineering

  • Speechify's Simba 3.2 model is ranked number one globally for quality, offering 10x more affordable pricing ($10 per million characters) compared to 11 Labs ($100) and OpenAI ($196).
  • Internal engineering culture prioritizes "10 good decisions per day" via agent orchestration, with engineers acting as QA and architects rather than just coders.
  • The company uses Cloud Code and Cursor as primary AI coding tools; token usage is budgeted tightly to prevent waste, with engineers incentivized by shipped production features rather than token consumption.
  • Weitzman rejects "token leaderboards" as incentives, favoring a "milk delivery" analogy where credit is only awarded when a feature is successfully shipped to production with no bugs and user adoption.
  • Hiring criteria have shifted from specific technical skills (e.g., framework expertise) to "technical aptitude" and raw intelligence (e.g., Math Olympiads, physics backgrounds), as existing knowledge can be taught rapidly in the AI era.
  • Weitzman asserts that hiring is easier for seed-stage companies today because founders can leverage agents to achieve high leverage, whereas growth-stage companies struggle to compete with the $15M+ compensation packages of OpenAI and Anthropic.
  • Speechify is actively developing synthetic data sets internally and partnering with data providers like Mercore to bypass data scarcity and accelerate model training cycles.

Market Outlook and Future Technology

  • Weitzman predicts that human-computer interfaces will shift primarily to voice over screens within five years, driven by wearables and improved voice LLMs.
  • He identifies pharmacology and biology as the most promising application of AI for the near future, citing personal projects using GPU clusters to sequence genomes for orphan diseases and design protein-binding molecules.
  • Weitzman believes the customer support market is saturated with 18 companies raising over $100M, with large enterprises building in-house solutions, making it a less attractive market than the core API space.
  • He compares the competitive dynamics of 11 Labs and Sierra, noting Brett Taylor's resume (Google Maps, Meta CTO, Salesforce Co-CEO) as a major asset for Sierra, while acknowledging 11 Labs' focus on voice-centric agent capabilities.
  • Weitzman views Meta and SpaceX as formidable long-term contenders, with Meta's data advantage tempered by regulatory constraints (GDPR) and SpaceX's advantage in manufacturing and energy storage (Tesla).
  • The founder emphasizes that "being the warrior" (hands-on product use) is critical for identifying user problems, with the CEO personally using Speechify's duplex models to generate feedback for the engineering team.