newsfilter.io
Interview, Fireside Chat

Simulating Humans at Scale: Simile's Joon Sung Park

  • Simile, founded by Jun, is an applied AI lab building simulations of human behavior and societies to guide decision-making in place of costly field tests.
  • The company's origins trace to the "Smallville" experiment at Stanford in 2023, which created a 25-agent town where generative agents with memory, planning, and reflection lived autonomous lives.
  • Smallville demonstrated emergent social phenomena, such as agents spontaneously organizing Valentine's Day parties and handling social dynamics like uninvited guests or failed invitations.
  • Prior to Smallville, Jun's team published "Social Simulacra" in 2022 using GPT-3 to simulate entire subreddits, revealing the potential for modeling long-term community interactions before instruction tuning was available.
  • The team concluded that current foundation models, often optimized for rational problem-solving, plateau in their ability to simulate human irrationality, subjective values, and diverse tastes.
  • Jun and co-founders Percy Liang and Michael Bernstein launched Simile to transition from breadth-focused academic research to depth-focused application and validation.
  • Simile validated its platform by simulating a U.S. population of 1,000 people, achieving an accuracy of 85% in predicting behaviors compared to self-reports, a key threshold for commercial viability.
  • The company operates as a SaaS platform where customers define target populations, which Simile grounds using data collected from partners like Gallup and specialized interviews.
  • Simile closes the "say-do gap" by collecting behavioral data (e.g., life stories) rather than relying solely on attitudinal data found in LLM training sets, creating a translational layer between what people say and do.
  • CVS is a key early customer; their lead for Human Insights uses Simile to map second-order market impacts of decisions, such as how an electric vehicle launch affects the perception of non-electric vehicles.
  • Unlike traditional polling, Simile's simulations allow for instant testing of thousands of concepts across diverse subpopulations and can simulate multi-agent interactions, such as earnings call dynamics.
  • The company addresses limitations of live A/B testing by offering scale beyond available ad platform populations and providing representative samples that users cannot easily recruit organically.
  • Simile's predictive evaluation uses Total Variation Distance (TVD) as a North Star metric, targeting a TVD of less than 0.15 for quantitative decision-making confidence.
  • The firm categorizes simulations into "convergent" (where small errors compound but outcomes remain stable, e.g., network hub formation) and "divergent" (where outcomes vary, requiring bootstrap resampling to calculate confidence intervals).
  • Simile is developing proprietary models to capture human diversity, viewing frontier models as a coordination "CPU" while their specialized agents act as the "GPU" of individual subpopulation viewpoints.
  • Data collection involves a reinforcement learning loop to optimize interview questions for maximum visibility in minimal time, supplemented by efficient surveys for factual data.
  • Jun predicts that perfect behavioral simulators could solve major societal questions, including macroeconomic modeling, climate change collective action, and detecting signals of democratic collapse.
  • The vision for the technology aligns with science fiction tropes: a society guided by two pillars of AGI and comprehensive social simulations to prevent negative externalities like harmful feed algorithms.
  • Simile aims to establish rigorous statistical standards and thresholds for social simulation, akin to the evolution of inferential statistics in natural sciences.
  • The company is currently open to leveraging customer in-house data to fine-tune specific models for unique populations, provided ethical and responsible usage is maintained.
  • Future capabilities include simulating long-term impacts over 5–10 year horizons, moving beyond immediate product testing to strategic forecasting for governments and corporations.