newsfilter.io
Interview

#30 - Dr Eva Vivalt on how little social science findings generalize from one study to another

  • The Y Combinator basic income study's shortest treatment arm is projected to conclude in three years, featuring a midline survey approximately two and a half years into the trial prior to data release.
  • Individuals receiving basic income are expected to reallocate reduced labor time toward productive activities like education or childcare rather than reducing labor supply irrationally, with the program likely altering life trajectories of young, poor participants more significantly than broader-targeted initiatives.
  • Conventional meat producers are predicted to attack clean meat by leveraging negative social information to exploit naturalistic heuristics, while advancing biotechnology will necessitate messaging campaigns to counter the "unnatural = bad" perception for new "unnatural" products.
  • Generalizing findings from development economics studies to other settings is expected to be poor, with effect size predictions often differing from true values by a median absolute amount of 99% or 0.18 standard deviations.
  • Results from small NGO studies often appear promising initially but fail to match outcomes when governments scale up programs, and within a single country, results from different locations rarely predict each other, indicating low generalizability even at the national level.
  • Studies with high underlying heterogeneity are predicted to provide the highest value for policymakers in specific unique contexts, whereas aggregating crowd wisdom or expert priors may yield better policy outcomes than conducting additional randomized controlled trials in certain scenarios.
  • Policymakers are expected to exhibit bias by updating positively on "good news" while resisting negative updates, neglect statistical variance in favor of point estimates from small studies, and requiring overwhelming evidence and wide data ranges to be compelled to acknowledge negative results.
  • Aggregating predictions from many individual experts is expected to yield surprisingly accurate priors despite individual inaccuracy, and researchers collecting priors before studies can demonstrate that findings were not known a priori to add value to null results.
  • False positive report probabilities in development economics are predicted to be smaller than initially feared, suggesting the field is in a better state than the broader replication crisis implies, with specification searching expected to be less prevalent in RCTs than non-RCT studies.
  • Non-RCT studies are becoming increasingly significant over time as the field grows more aware of specification searching, while the economic profession is not expected to become fully Bayesian soon, though there is growing acceptance of priors and prediction error over unbiased estimates.
  • While many social interventions show weak or null effects individually, meta-analyses of underpowered studies often reveal an average effect, and the benefit of picking the "best" intervention over a random choice may be smaller than expected as random selection retains a reasonable chance of high performance.
  • The true dispersion in cost-effectiveness across interventions is expected to remain very wide, likely following a log-normal or power law distribution driven by varying costs and welfare values, even if raw effect sizes appear normally distributed.
  • Future research is predicted to increasingly utilize observational data and graphical models, such as Bayesian networks, to narrow down causal mechanisms before deploying resource-intensive RCTs.
  • The academic job market for economics PhDs is characterized by high centralization and significant random noise, making success partly a "crapshoot" beyond work quality, with extremely high quantitative score thresholds acting as critical filters for top-tier programs.
  • Academic markets reward "hot" topics, which can lead to trade-offs where researchers focus on important but unimportant questions, while the number of unconducted impact evaluations introduces bias by favoring the evaluation of successful or feasible program instantiations.