newsfilter.io
Interview

#39 - Spencer Greenberg on the scientific approach to solving difficult everyday questions

Sparkwave Portfolio & Personal Updates

  • Spencer Greenberg manages four active portfolio companies at Sparkwave, having recruited CEOs for all and spun them out as separate entities.
    • Uplift: An automated depression care program run by Eddie Lou; currently in closed beta with plans for imminent public release.
    • Mind Ease: A software product designed to alleviate significant anxiety, managed by Peter Breitbart.
    • Clear Thinking: A tool for improving decision-making and reducing bias, led by Aurora Quinn Elmore.
    • Positively: A recruitment platform for social science research, run by Luke Freeman.
      • Currently utilizes Amazon Mechanical Turk as its first backend to layer features for researchers.
      • Plans to add new backends to secure representative national samples or specific demographic cohorts beyond Turk.
  • Clear Thinking Research & Products:
    • Conducted long-term studies on "happiness habits" using environmental triggers; Greenberg reported a 5% personal happiness increase by associating a daily trigger (checking social media) with gratitude.
    • Testing a methodology where pre-existing daily triggers are used to induce specific positive thoughts (e.g., gratitude) to create lasting habits.
    • The goal is to release this as a public tool if study results validate the efficacy of the trigger-based happiness technique.

Intrinsic vs. Instrumental Values Study

  • Methodology: Conducted a rigorous survey of Effective Altruists (EAs), partial EAs, and non-EAs across demographics (age, gender, politics) to identify intrinsic values.
    • Excluded participants who failed a quiz on defining "intrinsic value" (valuing something for its own sake, not its effects).
    • Participants were required to define the concept in their own words to ensure comprehension.
  • Key Demographic Findings:
    • Conservatives: Report higher intrinsic valuation of religion, retribution, and the preservation of existing values.
    • Liberals: Report higher intrinsic valuation of animal well-being, nature, and the happiness of strangers.
    • Females: More frequently report intrinsic values regarding kindness, caring, diversity, and human freedom.
    • Males: More frequently report intrinsic values regarding selfish interests, in-group interests, and stranger pleasure.
    • Older Adults: Value being trusted/cared for and general societal morality.
    • Younger Adults: Value animal lifespans, personal admiration, and the pleasure of those they know.
    • Effective Altruists: Show a distinct pattern of devaluing personal/in-group interests while strongly valuing the suffering and happiness of all conscious beings.
  • Universal Intrinsic Values:
    • 82% of non-EAs report "I love other people" as an intrinsic value; 71% value "beautiful things continuing to exist" even if unseen.
    • Universal values (those not tied to self or specific in-groups) serve as a potential foundation for global cooperation.
  • Practical Applications of Value Identification:
    • Avoiding Value Traps: Distinguishing intrinsic from instrumental values prevents pursuing careers or goals (e.g., high money) that do not actually deliver the desired intrinsic experience (e.g., autonomy).
    • Goal Factoring: Identifying the core intrinsic value allows for more efficient planning than pursuing intermediate goals (e.g., becoming a tenured professor) that were assumed to be the end-goal.
    • Social Guilt: Understanding that different groups hold different intrinsic values reduces feelings of alienation or "wrongness" when one's values diverge from family or community expectations.
    • Preventing Doublethink: Separating subjective intrinsic values from objective moral beliefs prevents self-deception where individuals claim to value only "acceptable" things (e.g., global welfare) while ignoring their genuine self-interest.
    • Future Vision: Building a desirable future requires balancing multiple intrinsic values to ensure broad appeal, avoiding "hedonic treadmill" scenarios like a world of only pleasure machines that ignore autonomy or beauty.

Overconfidence and Calibration Studies

  • Overconfidence Prediction Model: Greenberg identified five traits of a skill that predict whether people will be overconfident or underconfident regarding their performance relative to others.
    • High Self-Perceived Ability: Skills people believe they are good at tend to yield overconfidence.
    • Subjectivity: Skills viewed as matters of personal opinion (e.g., writing a novel) correlate with overconfidence.
    • Experience: Higher self-reported experience correlates with overconfidence.
    • Personality Connection: Skills viewed as reflective of character (e.g., making friends) correlate with overconfidence.
    • Perceived Difficulty: Skills viewed as very difficult correlate with underconfidence (e.g., running a marathon, knitting).
  • Empirical Findings:
    • People are generally overconfident across a broad range of skills but consistently underconfident in skills perceived as difficult, objective, or unrelated to personality.
    • A new tool is being developed on ClearThinking.org to predict overconfidence based on these five traits.
  • Calibration Training:
    • Greenberg created a quiz distinguishing common misconceptions from true facts to train users in calibration (predicting confidence intervals).
    • Results from the quiz highlighted the difficulty of verifying truth, citing the "spinach iron myth" which involved a chain of errors where the original error was a decimal shift, and the subsequent debunking was also a myth.
    • Calibration training demonstrates that people often give ranges that are too narrow (overprecision), though they can learn to improve this.

Bayesian Updating and Evidence Evaluation

  • The "Question of Evidence": A framework for estimating the strength of evidence using the Bayes Factor.
    • Formula: Posterior Odds = Prior Odds × Bayes Factor.
    • Bayes Factor Calculation: "How much more likely is this evidence if the hypothesis is true compared to if it is false?"
    • Ratios indicate strength: 3:1 (moderate), 30:1 (strong), 1:30 (strong against).
  • Common Cognitive Errors in Updating:
    • Ignoring the Comparative: Focusing only on the likelihood of evidence given a true hypothesis, ignoring the likelihood given a false one.
    • Dismissing Weak Evidence: Accumulating weak, contradictory evidence over time can shift beliefs significantly; ignoring "trickles" of evidence leads to stagnation or error.
    • Neglecting Priors: Failing to account for the base rate of a hypothesis (e.g., mistaking a friend in a random city for a known friend due to low prior probability).
  • Reliability of Priors and Intuition:
    • Trust in Intuition: Intuition is reliable only when there is high-frequency, low-noise feedback (e.g., a therapist reading patient cues). It fails in domains with delayed feedback (e.g., long-term therapy outcomes) or no feedback (e.g., philosophy, macroeconomics).
    • Reference Class Forecasting:
      • Start with an "outside view" (base rate of similar events) rather than an "inside view" (specific project details).
      • Bias-Variance Trade-off: Narrower reference classes reduce bias but increase variance due to smaller sample sizes; optimal forecasting balances these.
      • Example: Estimating startup success or project duration requires weighting historical failure rates (e.g., 90% of startups fail) against specific mitigating factors.
  • Handling Specific Evidence Types:
    • Models/Theories: Treat all models as wrong but useful; averaging predictions from multiple independent models reduces individual bias and noise.
    • Empirical Studies: Evaluate studies by estimating the Bayes Factor, considering sample size, methodology, and publication bias (e.g., "p-hacking").
    • Heuristics: Qualitative frameworks (e.g., "do I feel excited hiring this person?") can add value to predictive machines even if they resist quantification.
    • Sanity Checks: Compare estimates to broad constraints (e.g., a company's value cannot exceed global wealth) to detect citation errors or "citation worms."

Case Studies in Updating

  • Trump-Russia Collusion:
    • Prior: Set low (e.g., <10%) based on the rarity of presidents colluding with foreign powers.
    • Update: The meeting between Trump Jr. and a Kremlin contact increased the probability (Base Factor ~3-4:1), moving the odds from ~10% to ~30-40%.
    • Context: Updated by distinguishing between the plausibility of the act and the evidence for it.
  • US-China War:
    • Prior: Reference class suggests ~70% war likelihood during major power transitions (historical data).
    • Update: Adjusted downward due to the "regime change" of nuclear weapons (mutually assured destruction reduces war probability).
    • Warning: Avoid double-counting evidence (e.g., multiple tariff disputes may not be independent signals).
  • North Korea Nuclear Disarmament:
    • Prior: Low (~10%) based on few historical examples of nations voluntarily giving up nukes.
    • Update: Statements of intent to disarm were given a moderate update (3:1) but counterbalanced by the strategic incentive to keep weapons for security.
    • Outcome: Net probability likely remains low due to strong strategic incentives to retain weapons.
  • Theranos:
    • Analysis: The company's failure to release a product for years ("vaporware") shifted the hypothesis from "incompetent" to "fraudulent."
    • Evidence: Poorly written scientific papers suggested incompetence (sincere failure) rather than calculated fraud, as a fraudster might produce better-looking fake data.
    • Investment Signals: Lack of interest from knowledgeable Silicon Valley biotech investors served as a negative signal.
  • Power Posing:
    • Context: Original study claimed large effects on hormones and behavior; subsequent pre-registered replications largely failed to find these effects.
    • Greenberg's Analysis: Suggests a small, variable effect exists (n=1000 study found mood/power boost), but the original effect size was likely overstated.
    • Placebo Factor: Acknowledged that the "power" effect may be a placebo or self-fulfilling prophecy, which is still valuable if it improves mood.
  • Dietary Advice:
    • Conclusion: Strong evidence exists only for specific deficiencies (e.g., Vitamin C for scurvy, Vitamin D for older women).
    • General Advice: Most specific food claims lack high base factors; randomized trials are difficult to conduct, making causal links hard to prove.
    • Anecdotes: Rapid recovery from a long-term condition after a diet change can provide a high Bayes Factor for an individual, even if it doesn't generalize.