Panel, Conference Presentation
Big Data as Driver of Disruptive Technology Breakthroughs
Milken InstituteSumant Mandal, Andrew Appel, Vanessa Colella, Vikas Kapoor, David A. Steinberg, Andreas Weigend
Definitional Shifts in Big Data
- Big data is framed as a "mindset" rather than a volume metric, predicated on the assumption of free data and compute power to focus on asking high-quality questions.
- The evolution of data capability follows a progression: "Big Data" (massive repositories) $\rightarrow$ "Fast Data" (real-time processing via Moore's Law and falling storage costs) $\rightarrow$ "Smart Data" (decision-making in 1–3 milliseconds without human hypothesis).
- Future analysis aims to move from 2D data extraction to 3D extrapolation using machine learning, eliminating the need for traditional hypothesis-based analysis.
The Fragmentation Reality vs. Theoretical Riches
- The "Two Cities" Phenomenon: While consumer-facing applications (e.g., Google search) demonstrate extreme data sophistication (million results in 200ms), large enterprises often struggle with basic data integrity.
- Corporate Data Silos: Many large organizations suffer from exponential growth in applications and fragmented data storage, making "good data" harder to access than in previous decades.
- Integration Challenges: Companies face difficulties with "slowly changing dimensions" (e.g., changing card numbers or customer identities) and disparate naming conventions that hinder unified views.
- Scale of Legacy Debt: Typical large enterprises manage between 100 and 1,000 major legacy systems, with some telecom giants operating 28 separate data warehouses prior to consolidation.
Success Cases and Competitive Advantages
- Kroger: Leverages a 60-million-household loyalty file to deliver personalized weekly promotional books, achieving 51 consecutive quarters of same-store sales growth.
- Limitation: Success relies on in-store data; the next phase requires integrating mobile, credit card, and out-of-store behavior to achieve a true 360-degree consumer view.
- Walmart: Uses algorithms to adjust website prices approximately 10,000 times daily to maintain a competitive position at the 80th percentile of low pricing.
- Telecom Case Study: A major global telecom consolidated 28 disparate data warehouses to enable a unified repository, facilitating the sale of 1 million cell phones in four hours through precise customer clustering.
- Data Abstraction: One firm reduced 1 million data variables to 10,000 derivable abstractions, drastically reducing code complexity and operating expenses while creating a common language across the enterprise.
- Kroger: Leverages a 60-million-household loyalty file to deliver personalized weekly promotional books, achieving 51 consecutive quarters of same-store sales growth.
Data Governance, Privacy, and Policy
- Regulatory Divergence:
- Europe: Focuses on "left-tail" risk (worst-case scenarios), with strict laws like GDPR imposing fines up to 4% of global revenue for data loss.
- United States: Focuses on "upside" potential, though state-level attorney general interpretations create fragmentation.
- China: Utilizes data for a "social credit score" system where behaviors (e.g., jaywalking, counterfeiting) directly impact access to employment and education.
- Consumer Data Rights:
- Experts argue that consumers cannot effectively own or track all their data due to its intangible, ubiquitous nature.
- The prevailing view is that institutions must enforce governance, provenance, and usage principles, while consumers retain control over specific marketing permissions (e.g., opt-out of ads).
- Trust as a Critical Asset: The primary barrier to data utility is not technical but cultural; organizations must bridge the gap between technical data builders and business decision-makers.
- Regulatory Divergence:
Investment Thesis and Economic Value Pools
- Value Category 1: Companies with distinctive tech/algorithms that drive consumer acquisition and retention.
- Value Category 2: Automation platforms that reduce operational costs (e.g., saving millions of man-hours in supply chains).
- Value Category 3: "Data as a Service" entities that aggregate data with cross-company use rights (e.g., consortium models) rather than siloed single-client data.
- Value Category 4: AI and Machine Learning firms that automate decision-making workflows.
- Sector Opportunity: Healthcare is identified as the "biggest dark area" due to extreme data silos (medical records, genomic data), representing a massive opportunity for data aggregation and service.
Future Outlook and Strategic Recommendations
- Dis-economy of Scale: Larger, legacy companies face slower growth due to an inability to optimize quickly, creating opportunities for smaller, agile competitors who lack legacy baggage.
- Cloud Dynamics: Cloud solutions offer an opportunity to enforce data governance across fragmented systems but pose a risk of worsening silos if unattended.
- Core Priority: For corporations with revenue over $1 billion, data governance must be ranked as the highest corporate governance priority.
- Strategic Pivot: Successful organizations must move from optimizing existing processes to building new, end-to-end data infrastructure, treating data as a core utility rather than a byproduct.
- Global Trend: Data is becoming a linear differentiator where actionable data drives competitive success, particularly in emerging markets where data exchange is increasingly tied to service access (e.g., viewing ads for free mobile service).