newsfilter.io
Panel

Architecting the Agentic Enterprise | Traversal, Kong, Pigment & Twelve Labs | RAISE 2026

  • Panel Context & Focus

    • The discussion centers on transitioning AI agents from experimental POCs to reliable, autonomous enterprise production in 2026.
    • Moderator Kathy Gao (Sapphire Ventures, $11B AUM) frames the core question as "Can agents do anything useful reliably?" rather than "Can they do anything?"
    • Five founders represent distinct agentic domains:
      • Anish (Traversal): Agentic Site Reliability Engineer (SRE) for Fortune 500; focuses on autonomous system healing.
      • Carl (Kong): AI connectivity infrastructure for secure, scalable autonomous operations.
      • Eleanor Crespo (Pigment): AI business planning platform unifying teams and agents for governance.
      • Soyoung (12Labs): Multimodal video understanding models for media and security.
      • Shubo (Axiom Math): Mathematical superintelligence for formally proving enterprise system properties.
  • Pitfalls in POC-to-Production Transition

    • Infrastructure & Governance: Enterprises require trusted, governed backends to manage access, validation, and human-in-the-loop thresholds across thousands of users.
    • Auditability in High-Stakes Environments: Critical industries (finance, security) demand immediate audit trails to validate agent actions, as errors carry reputational and regulatory risks.
    • Security & Actionability: Security teams resist autonomous actions that could destabilize systems; moving from "recommendations" to "system healing" triggers significant friction.
    • Cost & Token Management: Daily production use causes significant token spend growth, forcing companies to calculate value and control usage permissions.
    • Political Tensions: Discrepancies exist between buyers (seeking labor efficiency) and users (fear of displacement or lack of control).
    • Non-Deterministic Failure Modes: Unlike traditional software (binary 500 errors), agents produce "200-type" errors where execution succeeds but outcomes are wrong, requiring full traceability to debug.
    • Accountability Gaps: CIOs and CTOs struggle to assign blame for agent errors (e.g., incorrect loan approvals), often stalling deployment on core value-driving tasks.
  • Strategies for Reliability & Debugging

    • Deterministic Layers: Using LLMs to generate code or mathematical proofs (deterministic outputs) rather than direct actions, enabling established debugging frameworks.
    • Comparative Evaluation: Replacing absolute scoring models with stable pairwise comparisons to "hill climb" model parameters effectively.
    • Human-in-the-Loop Optimization:
      • Low-Risk: Automate fully with full traceability (e.g., board recaps).
      • High-Risk: Maintain human oversight for strategic decisions (e.g., hiring plans, budgeting) where context is critical.
    • Monitoring Systems: Systems must flag the 5-10% of failure cases for human review to avoid "alert fatigue," as humans cannot effectively monitor 95% success rates continuously.
    • Context Management: Humans remain essential for providing the initial context required for agents to operate trustably.
  • Business Value & Adoption Statistics

    • Gartner Data: 40% of AI initiatives launched in the last six months are projected to fail to deliver business value by 2027.
    • McKinsey Data: While 88% of companies use AI, only 5% attribute value to their bottom line or EBITDA.
    • Edge vs. Core: Current AI value is concentrated in "edge" use cases; true ROI requires deployment in "core" business operations (e.g., loan processing).
    • Uber Case Study: Uber spent $3.4 billion in three months on AI with no realized value, highlighting the risk of unmonitored token spend.
    • Cost Trajectory: Token costs are predicted to drop to near electricity costs, shifting the model from experimental edge to core value.
  • Organizational Structure & Roles

    • Responsibility Shift: Digital natives assign agent responsibility to builders; enterprises shift responsibility to the business functions where value is created.
    • Emergent "Empiricists": A new class of users (power users) emerges, treating agent interaction as research to discover unanticipated use cases.
    • Transformation Roles: Temporary "AI transformation" roles (e.g., AI CFO) are currently accelerating adoption but are expected to disappear as workflows become standard.
    • Decentralized Ownership: Long-term ownership of AI systems is expected to migrate to end-user business functions closest to the workflow context.
  • Pricing & Value Capture Models

    • ROI Dual-Lens: Value must be measured via both productivity gains (e.g., 80% efficiency) and decision quality (optimizing margins/supply chain).
    • Model Stratification: CFOs will split spend between expensive "frontier" models for high-value tasks and cheaper open-weight models for routine work.
    • Outcome-Based Pricing: Shift away from raw token consumption toward fixed pricing or seat-based models guaranteeing specific business outcomes.
    • Utility Metrics: Industries like chip verification will price based on "intelligence per dollar" or speed of outcome (e.g., faster verification vs. cost of testing).
    • Commitment Models: Adoption of AWS-style large commitments drawn down across various use cases to manage costs.
    • Sovereign Infrastructure: Lasting value lies in proprietary data infrastructure optimized for agent querying (indexing, caching), as standard APIs are ill-suited for agent needs.