newsfilter.io
Conference Presentation, Fireside Chat, Interview

Uncharted: Big Data as a Lens on Human Culture

  • Project Scope and Data Sources

    • The analysis utilizes a dataset of approximately 5 million books (roughly 500 billion words), representing about 4% of all English books ever published, derived from the Google Books project.
    • The raw dataset initially contained 15 million scanned books, but rigorous cleaning removed texts with metadata errors or Optical Character Recognition (OCR) failures, leaving the final corpus.
    • To address copyright concerns preventing the release of full text, the researchers created an "n-gram" table tracking the frequency of word phrases (1-grams to 5-grams) annually over the last 500 years.
    • The dataset is publicly available for free download and has been integrated into Google's dictionary and search tools.
  • Cultural and Linguistic Trends

    • Language Evolution: Data confirms the historical shift of irregular verbs toward regular forms, such as the transition from "throve" to "thrived," which accelerated over the 19th and 20th centuries.
    • Collective Memory: Civilization exhibits a consistent trajectory of interest in specific years, peaking at the year itself and then decaying; however, the rate of this forgetting is accelerating rapidly in modern times.
    • Technological Adoption: The time required for inventions to reach peak prominence has shortened from an average of 65 years (for 1800–1840 inventions) to 26 years (for 1880–1920 inventions), indicating an accelerating pace of cultural adoption.
    • Fame Trajectories: Contemporary celebrities achieve peak fame significantly younger (late 20s) compared to the 19th-century average (age 37), and their popularity peaks higher but decays faster, with fame levels halving roughly every century.
  • Career Path Analysis (19th Century Data)

    • Actors: Achieved fame earliest (mid-20s) but generally did not reach the same peak popularity levels as authors.
    • Authors: Reached peak fame later (mid-30s) but achieved significantly higher fame levels than actors due to the prevalence of mass media for books.
    • Politicians: Achieved fame latest (late 50s or 60s) but ultimately reached the highest peak fame, particularly for heads of state.
    • Scientists: Reached fame levels similar to actors but at a much later age (70s), limiting the duration of their peak recognition.
    • Mathematicians: Despite the myth of early peak performance, mathematicians in this dataset did not see significant public recognition until their 70s and 80s.
  • Measurement of Censorship and Suppression

    • Nazi Germany (1933–1945): The dataset quantifies the effectiveness of Nazi censorship, showing a massive shift in discourse; 10% of names were suppressed more than five-fold, compared to less than 1% in English datasets.
    • Selective Suppression: Historians on Nazi blacklists saw a 9% decline in mentions, whereas writers of philosophy and religion on the same lists saw a 75% decline (factor of four), indicating targeted suppression of ideological content.
    • Automated Detection: A weighted-average model successfully predicted suppression and propaganda effects, distinguishing between expected mention rates and actual discourse during totalitarian regimes.
    • Modern China: Data reveals significant suppression of the 1989 Tiananmen Square events in Chinese books, where mentions vanish rapidly compared to English datasets, leading to the removal of specific charts from Chinese book publications.
  • Predictive Capabilities and "Cultural Inertia"

    • Future Projections: N-gram trends exhibit "cultural inertia," where words trending upward for 20 years tend to continue rising, suggesting predictability in cultural language shifts despite individual free will.
    • Satirical Forecasting: An XKCD comic extrapolated that if the growth of the word "sustainable" continues, it will appear once per sentence by 2061 and dominate all sentences by 2109.
    • AI and Singularity: The speaker remains skeptical of "strong AI" or a technological singularity in the near term, noting that computers remain poor at comprehending complex human language, with current successes (e.g., translation) relying on statistical associations rather than understanding.
  • Applications and Methodological Insights

    • Big Data Democratization: The speaker advocates for the use of accessible data (scraping, databases, visualizations) by students and researchers to identify novel patterns without requiring massive institutional resources.
    • Cross-Disciplinary Utility: Techniques used to analyze linguistic data are transferable to other fields, including genetics, where interactive visualization is increasingly vital for discovery.
    • Real-Time Analysis: The methodology is applicable to real-time tracking of social trends, stock market sentiment (e.g., movie buzz), and political polling to quantify public interest objectively.
    • Algorithmic Decision Making: The speaker acknowledges the trade-offs in using algorithms for hiring, noting that while imperfect, they are a practical necessity for filtering large candidate pools, though they struggle with quantifying individual nuance.