Fireside Chat, Interview
a16z Podcast | From Data Warehouses to Data Lakes
- Predictive analytics, machine learning, and AI will replace legacy reporting and business intelligence as the standard expectation for technology companies and investment banks, including financial institutions like JPMorgan Chase.
- Legacy data warehouses will transition toward data lakes consumed by modern visualization and algorithms, evolving into cloud formations hosted on providers like Microsoft Azure, AWS, or Google.
- Architecture will shift from the batch processing models of the 1990s to real-time streaming, where data arrives immediately rather than in monthly or overnight batches.
- A hybrid architecture will persist for the foreseeable future, integrating web and cloud applications with legacy systems such as COBOL mainframes.
- Data lakes will increasingly reside in the cloud for non-latency-sensitive data like marketing and web exhaust, while latency-constrained machine or factory floor data may remain on-premises.
- New technology stacks will layer web and cloud applications over multiple data lakes to enable "turbo mode" execution for predictive analytics.
- Self-service enterprise capabilities will grow, allowing business users to integrate systems via "build and buy" choices without relying on large programming teams.
- CIOs will delegate security and cloud strategy responsibilities to CTOs, focusing instead on scalability and strategic innovation to manage the optimal mix of on-premise and cloud computing.
- The number of enterprise applications or websites is projected to be at least 10 times higher than current large enterprise estimates, driven by decentralized business-level users.
- Data structures will fundamentally shift from rows and columns to document models like JSON and web exhaust, necessitating new "digital plumbing."
- Self-service integration will become an absolute requirement to move data from legacy IT environments into business-user accessibility.
- Predictive algorithms will model complex scenarios, such as forecasting gasoline prices for specific dates and locations, requiring continuous iteration of data and algorithms.
- Data lakes will store diverse formats including hierarchical, JSON, web exhaust, and machine data, as improving computing power extracts signal from noise over time.
- Real-time integration of trading exchanges and geo-signals will begin to influence sales and business processes similarly to real-time ad technology.
- Companies must manage a virtuous cycle where increased data production drives better analytics, which subsequently captures more data.
- Falling compute and hardware storage costs will create conditions for widespread adoption of data science and predictive analytics.
- The integration landscape will rely on APIs and JSON objects to connect diverse data sources for custom enterprise solutions.
- Organizations face a rate-limiting factor in realizing predictive value due to modern technology complexity and the need for intelligent architectural choices.
- The "millennial post-web generation" will continue to drive self-service capabilities as a fundamental enterprise software requirement.
- Tensions between chief marketing officers and CIOs are expected to disappear as CIOs become more business-oriented and marketing leaders become more technical.
- The concept of the data lake remains in evolution as cloud and big data worlds collide, with batch processing expected to disappear entirely in favor of continuous real-time ingestion.