newsfilter.io
Interview

#30 - Dr Eva Vivalt on how little social science findings generalize from one study to another

  • Y Combinator Basic Income Study

    • The randomized controlled trial (RCT) provides $1,000/month to ~1,000 low-income young individuals for 3–5 years, with a 2,000-person control group.
    • Total study cost is estimated at approximately $100 million over its duration.
    • Primary outcomes include time use, health, education, financial health, subjective well-being, and political/intergroup prejudice metrics.
    • Unlike negative income tax experiments of the 1970s, this study specifically tracks "what people do with their time" (e.g., education, childcare) rather than just labor supply effects.
    • Results will not be released until at least the end of the 3-year arm to avoid narrative distortion from interim data or subjects reacting to media coverage.
    • Motivations include potential technological displacement, utopian ideals of freedom from wage labor, and expanding the social safety net efficiently.
  • Generalizability of Impact Evaluations

    • Eva Vivald's research indicates that development intervention results do not generalize well; median absolute prediction error for effect sizes in new settings is ~99% in absolute terms or 0.18 standard deviations.
    • Effect sizes from smaller NGO-led pilot studies are significantly larger and more optimistic than those observed when the same programs are scaled up by governments.
    • Results within a single country do not reliably predict outcomes in other regions of that same country due to high local heterogeneity.
    • Statistical metric $I^2$ (unitless variance proportion) and $\tau^2$ (true inter-study variance) are used to quantify heterogeneity; for many development interventions, $\tau^2$ is high, indicating low generalizability.
    • Heterogeneity is reduced in health interventions (e.g., antibiotics, deworming) when controlling for baseline disease prevalence or incidence rates.
  • Decision Making and Value of Information

    • Studies on impact evaluations are most valuable for decision-making precisely in contexts where they have the lowest external validity (i.e., when the intervention is unique and existing knowledge cannot generalize).
    • Aggregating priors from a "wisdom of crowds" of bureaucrats or policymakers can sometimes yield better policy outcomes than running a single additional RCT.
    • Policymakers exhibit "optimism bias," updating readily on good news but resisting updates on bad news (neglecting negative results).
    • Decision-makers often neglect variance and confidence intervals, over-relying on point estimates from small, underpowered studies.
    • Providing more granular data (quantiles, ranges) helps mitigate updating biases by making negative results harder to ignore.
  • Research Credibility and Specification Search

    • False positive report probabilities in development economics are relatively low compared to other fields, largely due to large sample sizes in programs like conditional cash transfers.
    • "Specification searching" (fishing for significant results) is prevalent in non-RCT literature but less so in RCTs, which are easier to publish regardless of null results.
    • A "clumping" of p-values just above the 0.05 significance threshold (e.g., 1.97 vs. 1.95) suggests researchers are manipulating specifications to cross publication thresholds.
    • The "bias-variance tradeoff" is often ignored in economics; the field prioritizes unbiased estimates over precise predictions, despite prediction error being the ultimate metric for policy.
  • Consumer Acceptance of Clean Meat

    • Experiments to overcome the "naturalistic heuristic" (the belief that "unnatural" = "bad") found that prompting cognitive dissonance by highlighting other accepted "unnatural" goods (fermented foods, modified crops) was the most effective intervention.
    • Simply stating "unnatural is not bad" had a mild effect, whereas descriptive norms (telling people others are excited) were less effective.
    • Priming participants with negative social information significantly increased resistance to clean meat, suggesting potential for industry sabotage.
    • Knowledge of clean meat may positively shift broader ethical beliefs regarding animals and the environment.
  • Collecting Priors and Expert Prediction

    • Individual experts are often wildly inaccurate in predicting RCT outcomes, but aggregated expert priors are surprisingly accurate on average.
    • Systematically collecting priors (e.g., via bins or quantiles) before study completion provides a benchmark to identify truly surprising null or positive results.
    • Researchers are incentivized to collect priors to defend against hindsight bias ("I knew this all along") and to make null results more interesting.
    • The "gold standard" for elicitation is asking experts to assign weights to outcome bins to capture distribution uncertainty, though this is cognitively difficult for many.
  • Career and Academic Incentives

    • Incentive structures in economics academia discourage meta-analyses and data aggregation because prestige is tied to being the "first" on a topic rather than synthesizing existing evidence.
    • The academic job market is highly centralized (e.g., the annual conference interviews) and involves significant random noise in hiring outcomes.
    • Economics PhDs require strong quantitative skills (high GRE math scores), but success in research differs from success in coursework; many high-performing students fail to thrive in research.
    • Second-tier PhD programs can be viable for policy careers if integrated with institutions like the World Bank or near D.C.
  • Broader Implications for Evidence-Based Policy

    • While RCTs are criticized for poor generalizability, Vivald argues they remain the "only game in town," with the solution being more studies and better integration of observational data and priors.
    • Graphical models (Bayesian networks) using observational data could theoretically approximate causal estimates if the causal structure is trusted, potentially reducing reliance on expensive RCTs in some contexts.
    • The "Perils of Partial Attribution" argument suggests that a focus on rigor may discourage large-scale transformative policies that are harder to evaluate but offer higher expected value.
    • Effect size distributions in development appear normal (clumped), suggesting that the extreme dispersion in cost-effectiveness relies heavily on unreported cost data and varying welfare valuations, not just effect size differences.