Section Insights
Introduction to the Panel and Topic
What is the focus of today's panel discussion?
The panel will discuss architecting the agentic enterprise, emphasizing the reliability of AI agents in enterprise settings.
- The conversation has shifted from what AI agents can do to how reliably they can perform tasks in enterprises.
- Panelists come from diverse backgrounds, providing various perspectives on the topic.
- The focus is on the practical application of AI in enterprise environments.
Challenges in AI Deployment
What are the key considerations for deploying AI in enterprises?
Key considerations include auditability, security, control mechanisms, and understanding the budget dynamics between software and labor.
- Auditability is crucial in high-stress environments where AI makes decisions.
- Organizations must navigate the tension between labor costs and AI efficiencies.
- Understanding the political landscape of budget allocation is essential for successful AI deployment.
Traceability and Governance in AI
Why is traceability important in AI systems?
Traceability is vital to identify and fix errors in AI systems, as they can behave unpredictably compared to traditional software.
- AI systems require a different approach to debugging due to their non-linear decision-making processes.
- Establishing governance and traceability is essential for managing AI traffic effectively.
- Understanding the flow of AI actions helps in diagnosing issues and improving system reliability.
Deterministic vs. Non-Deterministic Outputs
How can organizations manage the outputs of AI systems?
Organizations should distinguish between deterministic and non-deterministic outputs, using human oversight for subjective decisions.
- AI systems can be designed to flag outputs based on confidence levels.
- Human involvement is necessary for handling low-confidence outputs.
- As AI models evolve, more tasks may become deterministic, reducing the need for human intervention.
Empowerment of Business Functions through AI
How is AI transforming roles within enterprises?
AI empowers end users within business functions, allowing them to build and manage their own AI solutions.
- Business functions are increasingly taking ownership of AI tools and workflows.
- The role of centralized IT may diminish as users become more capable of managing AI systems.
- Organizations must balance the need for centralized oversight with the agility of decentralized AI management.
Transcript
0:04 Hello testing. Good afternoon everyone. It's so great to be here. We have an amazing panel for you today talking about and can everyone hear me? Can you guys hear me in the back? It's good. Okay, great. I'm Kathy Gao. I'm one of the partners at a venture capital firm called Sapphire Ventures. We're based in Menllo Park. We manage 11 billion and we invest in the best enterprise AI companies. And today we're going to be talking about architecting the agentic enterprise.
0:36 a lot of the conversation over the past two years involving agents has been about what can agents do, right? Can they write code? Can they handle a customer conversation? Can they review a legal contract? But I think the more relevant question to be asking today in 2026 is not if agents can do anything useful, but can they do anything useful reliably within an enterprise setting? And our panelists today, I'll have them introduce themselves. Hey Anish. themselves briefly, but they come from very very different lenses. So this is going to be a great topic, but I'll pass it over to Carl to go first.
1:22 Thank you. Hello. It's great to meet you all. can you hear me as well? >> Good. So I am from Kong and other than what I would say and I'm a bit you know obviously biased have the best logo in the software industry. I love the gorilla I have to say. we are an AI connectivity company. So you know a lot of people will talk about reasoning and and the models and what they can and cannot do. We come from the angle of making that possible. So basically allowing enterprises to truly scale AI safely, securely to become autonomous, not just have a shadow AI system and I think a lot of topics we'll talk about today touches that. So that's the angle I'll come from. How do you really connect AI to to what you need below in your infrastructure and make that possible and make that work at scale?
2:08 >> Go ahead Elanor. >> Hi everyone. So Elanor Krespo, co-founder and CEO of Pigment. So Pigment is an AI business planning platform. So we unite business teams and agents on one governed scalable back end to really help you take faster better decision at scale. >> Hi everyone, I'm So Young. I'm one of the co-founders for 12 Labs. we build models and agents for multimmodal video understanding. so essentially we give LLM's eyes and ears to understand, comprehend, and utilize context of what's happening in video. we work a lot with enterprise and companies that are in the media, creative and advertising space as well as security. yeah.
2:51 >> I'm Shubo. thanks for having me here. I am the CTO and one of the founders of Axiom Math. We want to build mathematical super intelligence and the way it connects to the enterprise is we hope if we can prove very hard math problems we can formally prove properties of enterprise systems that we built today. >> Hello. hi everyone. It's nice to be here. Thank you for inviting me Kathy. My name is Anish. I'm the CEO and co-founder of Traversal where we're building an Agentic site reliability engineer. So think about whenever your system breaks, the act of troubleshooting it and figuring out how it broke and fixing it. We try to do that autonomously. very excited for the conversation because the the only kind of customers we work with are the Fortune 500 because that's where the scale of data and people is is large and the pain is the largest and so lots of thoughts around how do you actually deploy in the enterprise and get them to keep getting value and purchasing from you. and from previous life, I'm a professor at Columbia University. And so I've been in the world of AI for the last 10 years, way before it was cool. so yeah, >> amazing. Okay, to kick us off, raise your hand if you're currently running an agent at your company.
4:08 >> Okay, good number. Keep keep your hand up if you let the agent run by itself without a human in the loop. Okay, a couple a couple up here. so my first question to the panelists is a lot of agents look amazing in demo or in PC but the stakes are raised much higher when they actually go in production and oftent times there's a a gap between PC to production. Where do you see the biggest pitfalls? Maybe I'll just ask Eleanor first because you have such a high fidelity use case around enterprise planning.
4:42 >> Sure. So at Pigment so we serve very large companies. So you think Coca-Cola, Enthropic, Uni Lever and obviously what they want when they use agents is not just individual use cases. They really need to do that at scale across thousands of users in a trusted manner. So what we see and what we've seen really in the past year I would say and even in the past six months is that they've run a lot of experimentation but now they are thinking about how do we put that in production at scale and in order to do that what they need and this is what we provide with pigment is the right infrastructure. So they need an infrastructure that they can trust with the right governance where you know agents you really decide what they have access to, who validates what, where is the human in the loop or not for some that don't want the human in the loop.
5:27 how do you actually make it scalable? So how can you make these agents work across you know thousands of different use cases with thousands of different business users that have to interact with these agents. So you need this governed scalable backend that has the right level of trust. But not only that, it's not only about the technology, it's also about the change management. So obviously now what we see is that you know a lot of people have done a lot of experimentation themselves but now they need to change their processes at scale to really help the entire company embark the new journey and the new way of working. For us what we try to do really with all of our customers is autonomous planning where really we automate workflows end to end and we make your team 80% time more productive. So in order to do that you can imagine it requires a lot of change management for teams to know where they have to work and here I think is a responsibility of the CXO to actually make that happen.
6:19 >> I'd love to hear from you Anishh right when you're talking to customers they have you know SRRES and you're coming in with a new paradigm what are they worried about? So, interestingly, I'd say the worry is more around it's such a high it's such an important use case because when your system is broken, that's when you're losing money, there's reputational risk with the financial services, there's regulatory risk. And so, a lot of times the issue is about how can I validate that the answer you gave me is correct quickly and are you taking me down the wrong direction? And so that's one question that comes up a ton is like making sure that what we give them is auditable very quickly. And second, especially for now, we're helping actually take autonomous actions on behalf of the system. So it's no longer just a recommendation, but we're actually changing the system to heal it.
7:11 And then there's always the question of how do you make sure security asks the question a lot like how do you make sure you know bring my system down, right? By taking action. So that's the second question that comes up a lot. The third one that I think is that's showing up a lot now is that now people are starting to use the product for things outside just incidents. They're using it for their daily production use cases.
7:33 And so as you start using it for like daily production use cases, the amount of tokens people start spending on on traversal is is growing quite significantly. And so people thinking about how do I calculate the value of what traversal is doing? How do I actually give control so that the right people are using in the right ways? And so that's a third thing. And the fourth one I think is probably connected is the some ways the departure between the buyer and the and the user. The buyer in some ways wants to drive a lot of efficiencies in the company. A lot of it around obviously as you can imagine labor and the users have some sort of tension around that. So that's a fourth thing that one has to think about. So those those are four things. So one is like auditability in a high stress environment and if you're making taking actions on the system, how do you get security okay with it? Third is around how do you think about controls and and token use for for your system. And then the fourth is around the thinking about the politics in some ways around where the budget is coming from whether it's a software budget or labor budget and how buyers and users think about that. one of your points which is token maxing is another is a topic I want to come back to later but Carl any thoughts on the gap between PC to full deployment. Yeah, I think you know good points raised here. You raised the hands of people. I think the next question should be who are in production, right?
8:55 So we see a lot of demos, a lot of tests. And other than what was said here today that you need infrastructure for it to function, you need to govern your agents. You know, we talk about agents doing work while we sleep. And I think the important question is are they going to wake us up when we need to be woken up? Not the fact that we're just sleeping through the errors and issues that's coming up. But what I'm seeing a lot and Kong is quite fortunate. We work on both sides of the aspect. So our customers still on the digital native side and also the enterprises and all the big corporations that we spoke about here today. And one is deploying it faster than the other is of course the the the software natives. They're going a bit faster into production. They're they're running agents more and more.
9:34 The enterprises I talked to specifically the CIO the CTO's one of their main issues isn't just technology governance how to actually do it. they they tend to have an idea. Of course, you know, you need a you need a solid API infrastructure. You need to have all these things in place for it to function because if you don't, the agents are going to be dead on rival. It's not going to work. But if we are truly going to become an autonomous type of agentic world, which is what we want to become. We don't just want to, you know, human in the loop on the loop. We want to make sure we are we are there. But who is accountable for what the agent does? This is something that comes up a lot for the conversations we have. So, let's say an agent is starting to give out loans, right? Which is which is good, which is what you want them to do.
10:18 It's core business. And they start giving out loans outside the risk profile that they found something. They're doing it not perfectly correct to what you expect. Who's to blame? If someone is to blame, is it the person running the agent as in the the department where it functions? Is it the person that built the agent? And a lot of CIO CTOs I talk to, they have ability to go live or go in production, but they stay on the edge cases rather than the core cases where the value lies for Agentic AI because they don't know how to do it. They don't know how to place them into production and trust it. And then, you know, humans, we like control.
10:54 We like to know who did what. And I think we have a a path to go here in terms of how to figure that out. So you bring up a really interesting dynamic, right, which is and at a high level, we know that LLMs are and a lot of these agentic workflows are non-deterministic. How do you debug that, right? Like what if a loan agent goes off pie and does something that they shouldn't do? Like how do you figure that out? And maybe Shoubo, if you'd like to chime in because your company is focused on formal verification and you're dealing with math, very deterministic, but I'd love to hear your thoughts.
11:35 >> Yeah, I think perhaps one of the only ways you can make a probabilistic system have some guardrails is having it output something deterministic. And what is deterministic are things like math, things like code. And this is kind of where we see where LLMs are really starting to shine where you don't output an answer, you output something that will lead you to the answer and that intermediate step is the deterministic step. And once you have things that are in the deterministic step, then you can build you know lots of other tooling like as humans in software engineering we have built lots of deterministic things and lots of deterministic things are also there to protect. I mean we humans make mistakes all the time. so that kind of I think is the way to kind of do it rather than LLM's directly acting on something acting on something through a deterministic layer that is then easier to control easier to debug which we have been doing in code for a long time. So I think that sort of is the way.
12:44 So, Young, anything to add just from a model lab perspective on going from PC to production? >> Yeah. demos are wonderful. PC's are great. production is always a different story. the, non-deterministic nature of agents and AI systems is very interesting. We deal with that a lot because, we deal with video. And if you think about a video is actually the most undeterministic data there is because you have sound, you have visuals, it's all context and then you have the concept of time. So it's everything by nature is narrative and interpretation can be subjective even obviously for the human. a very common use case that we deal with like our customers build and you know we serve is how do you go from a library or existing library of videos whether you know whether that's for marketing maybe it's sports footage film and then how do you transform that into shorter form consumable content that can be scaled and personalized for distribution to viewers and in that if you think about the agent pipeline that goes in to serve that use case there may be a creative brief for a prompt that you start with.
13:56 Now you're you going back to your like for the content you own and then piecing the finding the moments that are relevant, piecing that together in a narrative and then you're supplementing that with a content generation pipeline where you may have like additional effects. You may have add scenes into it. That's a workflow. Now the human in the loop question is very interesting because what makes a good content is a naturally very human thing. And so that that has to come with human judgment and human taste. which is the difference huge difference maker between what is production ready and what is a great demo. and in that the best like the kind of what we're seeing evolve and that's very interesting is actually using the AI to identify the more deterministic things and automating that process so that the human is on the loop in the most critical things. So what the AI can now is used for in that pipeline is to do compliance checks after it's generated it to use understanding models to be able to do things like are there technical defects?
14:57 Is this does this realistically match the prompt that the user gave it? And so you're fully almost the in this system you're doing kind of almost automating a lot of the process from ideation to creation to compliance and checking of it. and then for the human then you're able to pass on the most contextual nature of what it is. and that's kind of where we see production ready workflows going for agents like in this in the in the video space. So you very interesting what what you said but just to add to that as well what we're seeing on the question is it it's not like software in the same where software fails it fails fairly big and you have an error code 500 and you know you have a crash somewhere and you understand that aspect of it agents don't fail big I mean they could have a big outcome a failure but in general sense they'll kind of go through the path of what they're doing and they'll execute it fairly well the the outcome may be wrong to what what you were expecting So more like 200 type error. So you need a full traceability from the beginning to the end because it will not take the same route twice or every time it will go different. It will add other agents. It will do different things to get the same outcome. So you really need to be able to trace it from where you start to where you finish.
16:15 Otherwise you can't actually fix the problem. You don't know what the problem is. You just know you have an error because you have an outcome that's wrong. But it's not like software in that sense where you have a big crash big 500 and yeah we can now debug it. You really have to have that light traceability and governance of your AI traffic. >> Yeah, I think one well first I completely agree with Shuba that the right imprint of LLM is code like the right way needs to interact with production system is write code that that is executable and then once you have code then all of the frameworks we've been developed for the last 30 years to think about troubleshooting of code apply. So that is indeed and that's also the efficient way of dealing with these things. I think if you have to go with LLMs, one framework that we use at Traversal a lot is that rather than trying to get an LLM, have an eval framework that gives you an absolute score, that typically is not very accurate and every time you run it, that score might change and it's subjective, but but doing comparative tests is actually quite stable. And so if you have two LLM, let's say parameters, models, setups, it's the it's pretty it's pretty stable in terms of you're trying to hill climb doing like a pair-wise comparison and doing many many like pair-wise comparisons. That's the right framework for hill climbing on different model parameters or setups versus saying here's an absolute score for a given configuration. And I think that leads to a much more efficient way of hill climbing. Yeah.
17:37 >> Anything else to add? so Young brought up a human in the loop and that's a really interesting topic because at some point do humans become the bottleneck? I mean if the agents are able to do a whole lot of work but a human still needs to ver verify something before it can move on. How do we build a system that allows for that? Anyone have any thoughts? >> Can start. Okay. So in my case clearly at this stage we still need humans in the loops because most of our use cases are on finance and supply chain and I'm pretty sure if some of you work in these domains you are all freaked out about the fact that the results might be wrong. So what we what we advise our customers is really to think about it through two lenses. The first one is to think about what are all of the lowrisk things that you can do where you can really remove the human in the loop. So what are all of the tasks you know whether it's like recommending I don't know like a recap for your B pack doing certain analysis that then you can check etc all of these loris things that you can really leave agents to run for you in parallel and given the fact that with everything you've said like agents today can be more traceable than a human you know so it's very very easy for you in these lowrisk task to actually completely check and audit everything that a human has done an agent has done sorry because everything is going to be logged extremely well. So this is completely low risk and I think on that you know you can really leave it to be done. Where we start seeing humans being a bottleneck is when an agent starts to take action and I think this is the next phase where we are going which is saying okay agents is going to recommend for instance what is going to be your hiring plan for next year or how you should think about your budget or your margins etc. And I think this is where it becomes more complex because obviously this is highly strategic and so this is where you really want to have someone that can really check before something is committed. The other thing where I think we still need a lot of people is around the context because for agents to be highly effective and highly trustable they need the right context and it's very hard for them to figure out where is the right context and obviously humans still need here at the first step to to bring that context.
19:53 But I would say when you separate that yes you will still have human in the loop at least for now even if the models are getting better and better but try to really do this this two lenses approach I think purely if we're talking about what this is going to be and not what it is today you know you may today be in the shadow AI part you may be in the scaled AI part that's majority of companies live and I also saw that you know some stats from Gartner that 40% of all the AI initiative started today or or or in the last 6 months is going to die in 2027 because they don't get any business value from what they're actually performing or doing. And I also saw from McKenzie 88% of all AI initiatives well 88% of companies are using AI is probably more but that's what they said but only 5% attest value to their bottom line or Ebel so it's really so you know human in >> you think companies are actually measuring this though cuz I was having a conversation earlier in the speaker lounge and this person was saying every company I'm talking to they're like okay we're spending a million dollars a month on anthropic but I have no idea what my employees are So do you think people are even at that level where they're actually measuring anything?
21:01 >> No, because they're at the edge. They're not at the core. So when we're talking about business plan or things that you mentioned, which is very important, but it's not core business for your company at the end of the day. So when I start producing agents that actually, you know, allow loans to happen or not happen and to what you talked about agents applying other agents into the workforce and we don't monitor that. we let them actually operate in the core of our business value. We will start to see that happen. But until we do that, we're not going to see the business value from any agentic kind of deployment because it's all going to be at the edge. It's all going to be safe and it's all going to be maybe human in the loop governing it which will stop what it's supposed to do. So you have those three things, human in the loop, human on the loop, human off the loop. And I think you're right, Elanor, off the loop. You have to have below the line not that you know critical use cases but on the loop you need to kind of set up a way where you have your agentic systems and it will it will work 95% of the time there'll be kind of the five to 1% failures at least what we see but you need a system to govern that first of all because if humans are in the loop or on the loop and look at that every day humans we are distracted I'm sure many of you are distracted already today about what we're saying or not saying but we can look at something day in and day out and not understand what we're looking at or get bored or at you know stop looking and then we miss it. So the system has to tell us these are the 5% you should look at then we will function otherwise we will just slow it down it will just fail it'll be like micromanaging people in your organization no one likes that so to become fully autonomous and move from those initial phases where there is no business value to actual business value you need again the governance layer you need to set it up properly so actually you can trust the agentic to do what it should do add business value >> yeah I think what's in the entire process and the chain what's indisputable is there's a at the end of the process there's a human that's accountable for the outcome and there's likely a business value or business purpose that is measurable in value and given those two things prior to those happening where the final output is made and you get to a stage where you can actually measure the value whether that's conversion you know sales conversion or engagement whatever that may super to determine like what can be automated. what what is something that you can trust the system to be able to do autonomously versus what is uniquely human judgment. as an like if we take a like using a creative example it actually takes so many prompts if you're using a generative model to create something that fills the need the creative need that a user wants. But in that process there are actually a lot of things that can be automated that's deeply that that's technical. and that may be things like does this you know is there very crude example is there a sixth finger is something that you know we are like we're all like familiar with or seen is there something unnatural that isn't you know tuned to like the physical world. and these are things that are much more deterministic. And if you can distinguish between what is that and have the system flag things with with confidence with confidence or scoring so that these things that are can be flagged with low confidence or high confidence and then pass on the things that are lower in confidence that are much more subjective and non-deterministic in nature. That's where the human in the loop can come in, human on the loop can come in and pass the final output. And I think a lot of systems can be designed in a similar way. but the models are al underlying models are also evolving very quickly which means if you as you evolve this system u much many more things can become in can fall into the deterministic realm. That's actually a great point because one underlying theme like you said so young is the models are getting better and as founders how do you think about what to invest in right like what is scaffolding that will still persist as the next generation of models come in versus what is temporary right or what is temporary scaffolding meaning like you could invest three months into something and then the next cloud model drops or open AI model drops and it's moot. How do you think about that?
25:27 >> So, I think that's a a question that everyone thinks about a ton. I can imagine from our perspective, if all of the innovation a company is doing is at the harness level, then I think you're in a you're going to be in trouble because the way you're interacting with data is primarily through the MCP service or the APIs or whatever. and all of the work in the harness is likely going to in some way shape or form get imbued into the model weights at some point. So I think then that that leads to the question of where will there be lasting value. In my opinion it's when you have to fundamentally change the data infrastructure level as well. And what that means is like for in our world of observability, the underlying APIs and MCP servers that data dog or Splunk or Graphana or Service Now give you are not designed for the way agents want to query data.
26:22 They're designed for things going to a dashboard, right? And when you have things going to a dashboard, you'll get really good refresh rates for the specific dashboards that you set up, but they're not designed for querying a pabyte of data quickly. And so we have to rethink all of the underlying data structures and indexing and caching so that it can be optimized for the way agents want to query data. And I think if you're doing that then the hardness is optimized for that new data layer that you that's proprietary to you. And and I think that's what we're seeing across lots of great companies of people let's say web search similar thing where if you try to you know have these agents query the web as it's designed today you're going to get rate limited in different ways. But if you kind of rethink a lot of the web infrastructure so that you can now agents can query in parallel a lot of different things then you're going to have a lasting mode right because that's something that's not going to be the frontier models are not going to be spending time on that on the data layer as well. So I think if you're building a harness that's optimized for proprietary data layer that's designed for the way agents want to query data I think that's at least the way we think about lasting value.
27:24 >> Yeah. >> Yeah. I I 200% agree with that. Like for us in the mathematical world we want to have the largest index or data structure whatever you want to call it of formalized knowledge because I think that's going to be the the moat and then everything sort of builds on that our models our agents so yeah rethinking the data infrastructure that is core to what you're solving and for us what we want to do is formalize all of world's knowledge that's where we think is the sort of mode.
28:03 So Carl, you brought this up earlier which is in the enterprise who is actually responsible for various agents right is it is it someone in the security you know department depending on the use case or does it actually live with the business I'm sure it's going to look very different at all companies but is there a highle philosophy or framework that people should be using to think about how they organizers.
28:33 >> It's a good question and I don't think I have the complete answer to exactly what they should do because I think it's still being debated of where the responsibility should lie. But what I'm seeing is a kind of 50/50 split between the people who built the agents will be responsible. I've seen that with more digital native type companies and on the enterprise side it tends to move towards the business side where the agents will actually operate and do tasks and and create value. That's where the responsibility is going to lie of whatever happens at the end of the day.
29:04 Should it be a core business critical use case that actually adds value? >> Are there going to be new roles that pop up like agent manager? What's your guys' prediction? >> I think so. I think we'll see new roles and new responsibilities for this to truly take shape and function in the future. And I'm actually already talking to CIOS that are telling me that we need to hire differently. we need to have probably a new type of roles for us to deploy this at scale because again projects is projects scale is different and we are at project stage in most cases today I think one very interesting thing we're finding with our users is that they kind of go into three groups so you have group one where they have a predetermined workflow that they always did and it's a bit of a pain in the ass and now they're using AI agents whatever to help them automate some part of their workflow let's say it's instant RCA whether it's postmortems, whatever.
29:57 Then you have but they're doing it for something they already knew existed. Then you have a second group where they just don't know how to use it and especially the engineers that are outsourced that don't really are not particularly familiar with that system. They're using it in ways that don't make any sense. and that's a group. And then you have a third group which which is very interesting where they are finding use cases for it that even we didn't predict. And so we call them like the power users where they're finding ways to use the system that you know we we couldn't have even imagined them using it. In some ways that was why cloud code was so magical when we started using it. People were using in ways that we didn't expect. And so I think as you think about roles at companies there's this group of people that I think are it's it's it's hard to think about ROI for them because if you just think about in terms of token spend it's like I don't really know where it's going. But then it's like it's like the VC world. you'll have some of these outsized outcomes in the way people use these things that you didn't predict.
30:52 And so they're almost becoming like these researchers. They're not researchers, but they're researchers for that specific function in the org. And they're finding use cases for you. And I think that's going to become an interesting like crew of people within a company. And how do you incentivize them, reward them for finding use cases for you and then bringing that to the rest of the organization is going to be very interesting because they're not researchers, but they are thinking like researchers. the empiricists in how they interact with with these systems. And so I think that's going to be a very interesting group of people.
31:24 >> maybe if I may add something that I see a lot and I don't know if it's the same in all functions but it's exactly like when we had the cloud revolution that took way longer where you had specialized roles that were created just for that and then disappeared. And here the problem is that you all have to be AIP extremely fast. So what I see the best enterprise company do today is to have transformation rules that are created just really to help you accelerate the transformation. So finance transformations, finance AI transformation etc. But I can tell you these roles will disappear. They will disappear because the real value will be in the business team. So what I expect is that in you know three five years max everybody thanks to all of what you said you know like it will be end to end workflows that anybody can run anyone can become an agent manager. I already hear some my customers telling me we are mini CFOs right now because we can manage agents. So I think this is really important to really help accelerate the transformation but I think after that the role are going to disappear.
32:24 >> I couldn't agree more. I think if you think about like who within the enterprise the AI really empowers is the end user which is the business function. We've seen a lot of different transformation or IT organizations that kind of collaborate with the business functions to understand criteria use case evaluation frameworks but at the end of the day the users who really own it and the lasting value of AI system for the where you build is going to be the context how do you evaluate it how do you build the workflow that's unique to your business function and those things are owned by the business function the the world of AI today even for the enterprise is that end users if given the time and resource they're able to build their own and so I think the ownership is going to go increasingly to the business functions to be able to build the things that they need and to be able to run it and they're the ones who are most like closest to the workflow itself. hence they'll be able to adapt and evolve it as they need. Now how do you manage and orchestrate across the entire organization that's you know that's to be seen. and that's probably a more centralized technology organization, >> but we have to figure out because very good points, but you know, Uber burned $3.4 billion in three months with AI spend.
33:41 >> Yes, I was just thinking about that. >> So, you know, and and they didn't get any real value out of it in terms of what they look at it and and that's 3.4 billion by the way, their entire AI spend in in about a quarter. So, now they have to look at it differently and, you know, cap user per how much they can spend with AI. But again, I go back to what I said before. I mean token cost my opinion maybe wrong maybe right we'll go down in the next couple years to electricity plus a little bit more it's not going to be what it is today and we will move from edge to core and then we'll really start see the business value of what this can actually bring us because today is just a lot of costs a lot of experimentation you have to do it a lot of powerpoints but unless you really put it in place properly and you govern it right you're not going to get the true value out of what it is Whimo will win over Uber just based on the fact that how it's built Autonomous driving versus drivers, right?
34:32 >> That was going to be my last question. I think we have time for one more topic. Carl already said his piece, but I'm going to put the panelists on the spot. One prediction about pricing AI value capture. You know, we talked a little bit about token maxing. I think that's one of the top topics that founders are asking me and my portfolio companies is like how should we price AI? you know, the first phase was like, hey, everyone should use AI as much as possible and now there's backlash on the cost and if it's actually positive ROI.
35:03 So, Eleanor, any predictions about that? Yeah, so as you can imagine, it's a huge topic for us because our primary buyer, the CFO is the first one that is freaking out right now and a lot of our customers are already like, how are we going to manage our budget and it's it's super hard, etc. So I see it through couple of ideas. First of all what we try to define ourself within our own use cases is really making sure that they find ROI on two levels. The first one is about productivity gains. So thinking you know if you make someone 80% more productive well that's that should be the price that you're willing to pay and of course like you do a discount on that because you don't want to pay the same.
35:44 And then the second that is the next level is the quality of the decision you can make with the platform. Because ultimately with pigment for instance the idea is that you're going to take better decision than a human because you're going to be able to optimize more your margin, optimize more your supply chain thanks to the platform. So the quality has more value than just productivity gains. But the problem now is much broader than that for CFOs. It's not just about like the value on on the use cases we provide. it is really more like companywide what they should do here and so I think what we see is a world that is splitting into two parts today so one part which is the frontier labs that are extremely expensive and they didn't have to decrease their price because the demand was so large and so fast that they didn't have to decrease their price and obviously the other companies that are pushing open weights models etc and say we're going to 10x decrease your cost and CFOs are thinking about that constantly. So the way we are thinking about it at at my company is making sure that we build to make sure that we can provide the infrastructure for people to be able to switch depending on what they want to do. Be able to obviously use frontier models where it's really necessary for like super high value task and being able to provide the simpler infrastructure where it's needed. I was hearing Clay from Sierra the other day saying you know you don't need Frontier Lab to actually do customer support for shoes. this is how you should be thinking about it, right?
37:07 Like so I I think that there will be more and more of this approach and what we also did and by the way this is for all of you. We are offering a token management app actually to really help everybody in a company maximize the way they can look at at their token management. So whether it's a business buyer, IT CFO, they all have to collaborate on that. So if you're interested, we offer it for free and we are actually partnering with fireworks.ai AI to to also make it like making sure that we combine the management and the fact that we can offer open weight models. So I invite you to look at that.
37:42 I think we'll see more fixed or seatbased kind of going back backwards >> really that's an interesting >> and the reason is because what we observe from like customers is they want outcome and here what what I mean by fixed is what's fixed for a reasonable a passable outcome that gets them to where they need and if they need to go beyond that that can be an additional consumption but how do they have the assurance to know if and predict for me to be able to create a certain outcome and make a certain level of impact for what they're trying to do? to be able to predict that and that can be bas and that can be many different units and depends really on the product and the use case, but it could be based on the project, it could be based on the output. it could also be based on the business outcome. but to have that predictability and it's you know depends on the product and the use case to be able to determine what is that right level of deliverable that the customer is comfortable with. but I think we're seeing more movement towards wanting that level of predictability especially as model capabilities are coming up. and now it's up to like okay now we know what good means. Now how do we make this efficient? How do we make this predictable? and then the level of how do you go beyond that can you know we might go back to like that that is a level of how do we do the extra mile but that's not baseline any super last minute 10 seconds >> yeah some of the customers that we have been talking to I think we are soon going to start thinking about utility per intelligence and intelligence per dollar like for example chip verification is one of the areas that we're starting to work in and the questions there is can I ship my chips faster if I can verify them and for us it's like how much does it cost to verify a chip if it takes more than actually just testing it then it's no utility whatsoever so it's those kind of discussions that's so it's more kind of outcome based and how do you get to the outcome faster and cheaper >> I'll say quickly I think it's going to look like AWS commitments so people will make large commitments to a company and you draw it down through different use cases. Yeah, that's basically what's going to happen.
40:06 >> Okay, thank you everyone.
Summary
- The conversation centered on the reliability of AI agents in enterprise environments, questioning their ability to perform tasks consistently.
- Panelists discussed the importance of having a governed, scalable backend to manage AI agents effectively across various business units.
- Change management is crucial for organizations to adapt to AI, with CXOs needing to lead the transformation process.
- Concerns around accountability for AI decisions were raised, particularly in high-stakes scenarios like financial services.
- The need for traceability and auditing of AI actions was emphasized to ensure compliance and security.
- Panelists noted that human oversight remains necessary, especially for complex decision-making and context-sensitive tasks.
- Predictions about the future of AI pricing suggest a shift towards fixed or outcome-based models, with an emphasis on predictability and utility.
- The emergence of new roles, such as "agent managers," was discussed as organizations adapt to the integration of AI into their workflows.
Questions Answered
What is the focus of today's panel discussion?
The panel will discuss architecting the agentic enterprise, emphasizing the reliability of AI agents in enterprise settings.
What are the key considerations for deploying AI in enterprises?
Key considerations include auditability, security, control mechanisms, and understanding the budget dynamics between software and labor.
Why is traceability important in AI systems?
Traceability is vital to identify and fix errors in AI systems, as they can behave unpredictably compared to traditional software.
How can organizations manage the outputs of AI systems?
Organizations should distinguish between deterministic and non-deterministic outputs, using human oversight for subjective decisions.
How is AI transforming roles within enterprises?
AI empowers end users within business functions, allowing them to build and manage their own AI solutions.