transcribe

AI Agents Are Creating A Data Explosion. Here's What To Do About It. — With Clint Sharp

Alex Kantrowitz · 37m · transcribed Aug 2026
More from Alex Kantrowitz Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Understanding AI and Telemetry

What is telemetry and why is it important for companies?

Telemetry refers to the data emitted from systems that run a business, which is crucial for understanding application performance and user experience. As businesses increasingly rely on software, the amount of telemetry data generated is growing, necessitating effective management.

  • Telemetry data is essential for observability and performance tracking.
  • The shift from phone interactions to app-based services has increased the volume of emitted data.
  • Companies must adapt to manage the growing amount of telemetry effectively.
# 7:34

The Challenge of Managing AI Telemetry

How should companies prepare for the influx of telemetry data from AI products?

Companies are often unprepared for the massive amounts of telemetry data generated by deploying AI tools across many desktops. This data is critical for security and operational insights, but many organizations lack a strategy to manage it.

  • The deployment of AI tools can lead to unexpected data management challenges.
  • Security teams need access to telemetry data for effective incident response.
  • Companies must develop plans to handle the telemetry generated by AI applications.
# 15:08

Future Use of AI Data

What potential does the explosion of AI-generated data hold for enterprises?

While enterprises are evolving slowly, the data from AI agents can help improve user experience and detect abnormal behavior. However, many companies are still figuring out how to effectively utilize AI, indicating a gap in readiness to leverage this data.

  • AI-generated data can enhance user experience and security monitoring.
  • Many enterprises are years away from fully utilizing AI data for improvements.
  • Startups may lead the way in leveraging AI data more effectively than traditional enterprises.
# 22:43

Security Implications of AI Data

How can companies protect themselves from AI-related security threats?

To defend against potential AI-related threats, companies need access to the same data as attackers. This includes understanding vulnerabilities exploited by AI systems, which requires proactive measures and collaboration among security teams.

  • Defensive strategies must match the capabilities of potential AI attackers.
  • Collaboration among security teams is essential for identifying vulnerabilities.
  • Government intervention and fear-mongering can complicate the landscape of AI security.
# 30:17

Cost Management in Data Processing

What are the financial implications of managing telemetry data for companies?

Managing telemetry data is costly, and companies often view it as a necessary expense rather than a revenue-generating activity. As data volumes grow, organizations must find cost-effective solutions to process and analyze this data without significantly increasing their budgets.

  • Telemetry data management is a cost center for most organizations.
  • Companies need to innovate to keep data processing costs manageable.
  • Future strategies must focus on efficiency and performance in data handling.

Transcript

0:00 How safe are AI products really? And how should companies manage the exploding amount of data they're now working with? Let's find out from Cribl co-founder and CEO Clint Sharp, who joins us in a conversation brought to you by Cribl. Clint, great to see you. Welcome to the show. >> Alex, it's great to be here. >> So, just to give folks a background of what Cribl is, it's a 9-year-old company, 1,200 employees, 350 million ARR. And what Cribl does effectively is help people make sense or help companies really make sense of the exploding amount of data that exists today and that seems to only be growing as AI agents start to proliferate around the web.

0:38 tell us a little bit more about what you know, telemetry is and you can maybe explain a little bit more beyond what I just shared. >> Yeah, for sure. So, it's really the the data emitted from the systems that run the business. So, if we are thinking broadly about, you know, what is it? When I interact with software, I there's a bunch of exhaust that's emitted. So, for every order that I, you know, complete in an online shopping cart, I requested hundreds of images and went to dozens of pages. And so, there is data emitted about all of that and I need that data to do a couple of things.

1:16 One is what is the performance of that application for me? We call that observability. And if I want to get down to like, hey, what was Alex's individual experience? And if you think about where we're at in the world today, you know, I'm a little bit older. I remember interacting with airlines on the phone. Used to call them, talk to a customer service rep. Everything was done, you know, all over the phone. Now, like I never interact with anybody over the phone. Everything's done on an app. So, I log into my United app. I do all of my my work there. What if that United app fails me in particular? Just me. Like, how do they have the ability to dive through a bunch of And this is telemetry data to see, hey, what was Clint's experience on this particular time? What problems might he have had with that with that application?

2:03 Or, on the other hand, maybe I'm trying to understand the security of my organization. I need a lot of high-fidelity information about everything everyone is doing in order to see whether some malicious actor is going off and doing what's called lateral movement, moving between systems, you know, investigating, you know, how they might have exposed factor, like what did they get access to? I want to have some comfort after some type of security breach that, you know, I have a way of going in and looking through all of this data to understand what happened. And that's telemetry data, and it's massive. And our customers our largest customers move petabytes of information daily. Whereas, like a data warehouse as an example, like that might be a petabyte data warehouse is a huge data warehouse. Like that's for for like a Snowflake or a Databricks, like that's a big data warehouse. And we've got people who are moving that every day, and they're keeping years worth of data. And so, this is very, very big, and it's growing very quickly.

3:02 >> So, there's two aspects to this that I actually want to speak with you about. And I'm very glad we're talking, cuz this is very pertinent right now. the first is in terms of how AI models learn from that exhaust, as you say. and then, what all that data actually does for companies' ability to compete, to deliver more services, and and, you know, to continue to be more functional in what they do. Let's talk about what the models are going to learn from this exhaust. So, I'm sure you've seen, a couple weeks ago, Satya Nadella wrote this piece. He published it on X. It's called the reverse information paradox. And basically what he said is, when you buy an AI solution from a foundational model company, it doesn't work like a traditional product, because they're actually learning from your use of that product.

3:52 and so, this is what he says. He says, "This requires more than data protection. Model models learn from exhaust that the prompts that people write, the tools agents use, and especially the corrections people make when the model is wrong. Every correction is distilled into institutional know-how. It's the the kind of knowledge a competitor could never buy, and the kind that leaks almost imperceptibly, trace by trace, correction by correction, eval by eval." So, I think we'll get to the good uses of what you would use this exhaust for, but I we've been talking about this on the show recently about like what these models might use the exhaust of our interactions with their products for, whether that's disintermediating our own products or building new things that we don't get a cut in.

4:38 as someone who works in this day-to-day, how right is Satya Nadella in picking this out as a problem? >> Well, I mean, it's particularly a concern for me because, you know, I I I have a near 100% belief that the that these companies will eventually want to compete with my business. and so, you know, I need them. I need their models. I need the intelligence that they're selling. But, I'm I have some very real concerns that, you know, within the next 6, 12, 18 months, they're going to have telemetry products, security products that are going to be directly competitive with mine. And I think most of the enterprises I'm talking to have similar concerns. Like the you know, they're starting to build you know, large forward-deployed engineer forces that are, you know, presumably out to help these enterprises build great applications, but also are they helping their competitors with all of that knowledge that they are that they are accumulating. But, that is exactly the type of data that we that we deal with. And one of the interesting sort of things that's happened over the last couple months is a lot of my customers are now asking for AI observability.

5:45 They want to look at this trace data. They want to know what's this costing me. They want to know what tool calls are these agents making. And that from a security perspective, like I need like rich auditability about these autonomous intelligences going out and and doing things on on my behalf. And it's a problem that did not exist 6 months ago. They had no idea they needed AI observability. It literally just appeared overnight.

6:15 and that's part of what we're helping people solve. The is is getting the ability to work with all of this data and ask and answer questions of all of this data that we probably didn't plan in advance. I didn't know I needed AI observability 6 months ago. How am I going to give you the ability to make sense of all of this? >> Right. It's like you when you see an AI work, it's calling all these tools, visiting these websites, all in service of solving your problem, but that's often fairly opaque.

6:45 >> It's It's very opaque and then it's also I I'm not in the loop. And if I put myself in the loop, well, then now I'm destroying the efficacy of what I'm asking this thing to do. I want it to be autonomous. I want it to go do things on my behalf. But how do I gain confidence that those are good and wholesome and right things. And you know, you see posts on X all the time, you know, model goes awry, accidentally deletes data, takes the wrong action. I mean, these are very real consequences of having autonomous intelligence do things on your behalf. And so that that exhaust that Sacha's talking about, that's exactly the type of data that that we are helping people work with. And you can see it already with just, you know, I have a customer, you know, a normal run-of-the-mill brand that you would know.

7:42 Not a company that you would think is like an innovative company. Like it's multiple hundreds of years old. They're rolling out Claude to 50,000 desktops. 50,000 that like that is insane. And every person who's using that is now going to be emitting a massive amount of telemetry that they had no plan for. They didn't wake up at the beginning of the year and say like I need to have a plan to ingest telemetry from 50,000 desktops.

8:08 But still the security department for sure is very very concerned. Like I need all this data because if I need to go back in time and see what did it do, I can't ask questions of data I don't have. >> Right. So there's there's I think we're talking about two issues here, right? So the first one is is are these AI companies going to take this exhaust data and sort of build products that compete with your products or other companies' products? And then it's like when you when you want to have like security review, like what do you do with all this exhaust that the AI is leaving? So let's go one by one.

8:40 I I suppose like the right question ask you on the the AI companies competing with others is how valuable is this exhaust data? Because that's sort of like Satya when Satya talked about it that the your use of AI model leaves exhaust and then the AI model companies can come build competing products. I just want to ask you like how true is that? How like is this data actually valuable enough for an Anthropic or an Open AI to see people using their models in certain ways and then being able to go out and basically go up market and build another product.

9:18 >> So my co-founder likes to say the value of this data is essentially zero until suddenly it is not. and so it's a lot of like individually looking at one user's traces essentially worthless unless I have some security problem I need to go urgently look at it, but normally it's worth nothing. But in aggregate, it's incredibly valuable because the being able and and it's also valuable it sliced different ways, right? So, there is value in saying, "Okay, hey, watching every developer, what questions they ask, what code gets generated, whether that code is good or bad, will allow me to reinforcement learning on the on these models to make them better for everyone."

10:05 Then there's also, okay, well, there's also organizational specific value, which is the the model, you know, working inside of an organization. We believe we will be doing reinforcement learning for specific organizations to tune the models to their specific environments. and so, the and there's also there's also value from a security perspective as well, both individually and in aggregate, understanding what is good and and what is bad. So, it's valuable, but it's also incredibly expensive.

10:41 This This data is emitted at very, very high velocity and very large volumes, and so, you know, we're having to build new technologies for how to process this cost-effectively so that we can actually use it because if it it's probably not worth an to an organization, this data is worth a few million dollars. It's probably not worth 10 million dollars. It's probably not worth 100 million dollars. But if I just collect and store all of it naively without doing anything smart with it and without processing it and summarizing it, the cost will be tens of millions of dollars, and then the then the business case doesn't work out. So, we do have to think very carefully about the business cases and how we go about getting value out of this data.

11:24 >> Right. So, this I I'm sure you remember the CNBC interview with Alex Karp where he was basically saying, "If you're using an AI product, they're charging you so cheap in order to sort of obtain this exhaust data and and come and and eventually take your bread and butter. I'm trying to to assess like so it sounds like that is a possibility but not as easy as Karp made it out to be. >> Yeah, I do I think they're there's I find that a little bit hard to believe. I mean, I think the they're certainly OpenAI and Anthropic in particular they're not cheap.

12:08 Not by comparison to what we're looking at with OpenWeights models etc. So, I don't believe that they are possibly subsidizing you know, token their premium tokens in order to get this this reinforcement learning data. I think there probably are I think it's more likely to come from like a Cursor or or somebody else where they're actually you know, definitively building their own models based off of based off of this data and those data sets are valuable. I mean, SpaceX bought them for $60 billion. A lot of it has to do with the with with that with that data set because Composer 2.5 is more cost-effective than than Opus 5 or Fable 5.

12:48 and they've been able to make it more cost-effective because of the data set that they have. >> I see. Okay, so now let's talk about how how this can be valuable to to companies. So, there is again, all this data that's submitted as you go about doing things within software. first and foremost, what is the amount of data increase that we're seeing as agents start to get off the ground? I think Greg Brockman was telling me a couple months ago that like only 5 to 10 million people use agents.

13:20 but I'm curious to hear your perspective on like are you seeing it now in terms of the increase in the amount of data that's being generated and where do you expect this to go as, you know, these agents aren't used by like five or 10 or 15 million people, but by a billion people? >> Yeah, so I believe it's felt, but I don't think we can measure it yet. So, the the general growth of data as measured by IDC is about a 30% compound growth rate. And that that's been that's gone up from let's call it 25 27% over the last couple of years.

13:51 But the to your point, agents are only in the hands of a handful of people today. it's not only in the data that's output by by using the agents themselves. It is also the other exhaust as the agents move at machine speed as opposed to human speed. So, if I now have agents that can work 10 20 50 times faster than humans, all the things that they are replicating that humans used to do, now also amplifies.

14:24 So, if that agent is browsing a website, and I used to have, you know, let's say I'm an enterprise, and I used to have 10,000 internal users of this application, but now I have agents which are, you know, exploring multiple hypotheses, and now it start to look like 30,000 or 50,000 users. Well, now I just had a an increase of a factor of three or five on all of this other exhaust in all my other systems that that are emitting this this data.

14:50 That being said, it's still too early to be able to point to like an industry example of like, well, we know it's going to be blah. So, logically, we're feeling it. We know it's happening. you can kind of get some samples and some observations of AI native organizations generate data at a significantly more rapid pace, but I don't think that we can give you a firm number, whereas I think in a year or two years, we'll have a much firmer grasp.

15:14 >> Do you have any thoughts? I mean, I guess you you've already touched on this with having companies monitor their their AI orchestration data. Do have any thoughts about like what an explosion in data coming from agents might enable companies to do in terms of using that data to improve what they do? >> I mean it it is Enterprises are are evolve slowly, right? So the it's largely the it's it's largely the same but bigger. The problem that that we need to solve is I are my users or in this case maybe now it's my agents getting a good experience and are they generally behaving well or how do I detect, you know, abnormal behavior that is relevant from from a from a security perspective. And then the third new one is can I use this information in order to improve these models.

16:16 And kind of where I see most enterprises at, I think they are likely years away from being able to do anything that that's sophisticated. I think that, you know, startups and and people at the technology frontier, companies like mine, like we will be able to do that and we will do that and we will do that some on behalf of of enterprises. I but many of the enterprises I'm working with like they're not even sure yet how to start consuming AI.

16:46 And so like the future is here, it's just very unevenly distributed. >> Yep. And from a security standpoint, I mean do you think this data will be more and more useful as like for instance if we have more instances of what happened with OpenAI and Hugging Face where the model effectively went rogue, hacked into Hugging Face, and now we know that they've tried they had like 17,000 different actions that they took within that database. I imagine that like as companies handle security, some of the data that you're looking at will be very helpful in terms of sort of tracing and and testing whether the AI agents are actually secure as as you let them loose.

17:23 >> And this is already happening. and even slow-moving enterprises are going to have to deal with the fact that you know, attackers have access to frontier level intelligence that that moves very very rapidly. And so, I think, you know, we obviously went through the whole sort of mythos scare tactic, but it's real and it doesn't need to be mythos. Like the what I have available in K3 or GLM 52 is better like is probably more skilled than the average attacker was before.

17:57 And they're going to move super super fast. and they and they already are, which is why my customers, you know, many of them are looking towards how how do I detect malicious activity? How do I get that time down to being measured in milliseconds to seconds? Because if because I have to be able to respond in an automated way because the attackers have are attacking in an entirely automated way. And then there's a whole set of ramifications from there that are also difficult and scary, which is okay, responding inaccurately or wrongly or one simple hallucination. So, let's let's say that the response is I detect some malicious activity and I'm going to attempt to isolate that activity on my network.

18:47 And I want that I want that malicious activity to stop. But some model hallucinates how wide of a block it should put in and then suddenly the the network ends up in a split brain or completely down. Like these are all very real concerns. So, the on one hand, I have a concern that there may be an attacker that can move so fast that that I can't possibly detect or stop them if I have to put a human in the loop.

19:14 And then operationally, I'm also terrified that I may have autonomous intelligence making production changes to my systems that may accidentally operationally cause me a a problem in the name of security. So, it it's going to be a a very interesting next couple of years while we figure out how to both defend and operationalize the defense of of these types of attacks. >> Hey Clint, you told me that many of your customers are now just trying to figure out how to consume AI. And in our conversations in the past, you've mentioned that like you get a groundswell of support to use AI from let's say C-suite and people on the ground, but then you go to let's say legal or finance, and they won't allow it to go forward.

20:00 >> they're terrified. >> Is this part of it? Is it Is it Is it because you just kind of give up control or what exactly is holding up the rollout here? >> I think it I mean, fear. Almost always. >> Okay. >> Yes. and the the fears are will my proprietary data, like we were talking about earlier, ends up end up in the hands of of these model providers? Will that be used to train models in the the future?

20:34 you know, what if I I'm using, you know, AI operationally as a site reliability engineer or as a SecOps engineer, what if it makes the wrong changes, just like I just mentioned, you know, what what hap- how do I control operationally, you know, what's what's going on? so, many many fears. And more than that, the things that move slower are more things like procurement, legal, right? Like the So, when we go sign an agreement, we have AI terms in there, and then then they come back with redlines, right? Like like no, no, for for me if you want to sell to my organization, these are the terms that that you must accept, or you can have no AI terms in there.

21:16 So, you'll literally have, you know, a chief information security officer on a podcast talking about what they're doing to protect themselves against mythos, and I can tell you that I'm selling to them right now, and you're telling me you won't use any of my AI technology, which will help you defend yourself against the these sort of things. It's getting better, and then that's a temporary I mean, in 5 years we'll have forgotten that that this was this was even a thing. But, it is definitely an impediment because the the thing that is 100% absolutely true and immutable is that the attackers have no such limitations.

21:53 >> Yeah, in the the OpenAI Hugging Face episode, I think one of the things that Hugging Face mentioned was that they were trying to figure out what was going on, who was hacking them, and they were getting a bunch of refusals from the safety guardrails in the models, and the attackers did not obviously didn't have to deal with those refusals. >> The best thing for everyone is that everyone has access to the best stuff. Like, the we we we cannot you know, we cannot guardrail our way into safety.

22:24 Because they're because as we've seen, you know, Chimera 3 being Fable 5 level trailing it by only a handful of months, the you can't will your way into this technology not existing. I can't just, you know, stick my head in the ground and pretend that it doesn't exist. I need to rapidly secure my enterprise and minimize the ability for offensive attackers using this technology to be effective. I need to have the same access as a defender to go patch myself so that I can secure myself from from from this technology. And I think that we're absolutely that that did happen. My CISO talked to as group of CISOs went and talked to the the defending team at Hugging Face and this absolutely happened.

23:14 They had to go to open-source Chinese models in order to look for the vulnerabilities that were exploited by by OpenAI's jailbreak. >> >> And that to me is just baffling. The only reason this is happening is because you know, because of government interventionism, which is also happening because of the fear-mongering that Anthropic has done over the course of the last 6 to 12 months trying to convince the world like, "Look, if you run around everywhere telling people that you invented a nuclear weapon, I think they're probably going to get scared and they're probably going to try to prevent you from giving anybody else the technology." But the reality is is that this same weapon that's being invented is being invented by eight to 10 different frontier organizations all at the same time. And so you can't you can't non-proliferate when that many people are building the same thing at the same time. So thusly, we just have to move forward with everybody having access to the best.

24:10 >> All right, Clif. Now, let me ask you about Anthropic in particular. we've had a debate on the show for a while about whether the myth of stuff was marketing or whether it was like a real area for concern. And you know, if these models are capable of like these long horizon plans and they can bug find and then they can go find zero days and sort of hack into entities like Hugging Face, what would have been like the right approach here in terms of sounding the alarm, guardrails, oh you know, etc.

24:43 etc. because you know, one of the things that I've sort of you know, brought up here is if you know, the capabilities seem real enough that I I wonder exactly like you know, do you do you just release him to everybody or where do you draw the line and how much of this mythos situation do you think was sort of irrational hype versus based in some reality? >> Well, I mean there's all kinds of rumors swirling like the really it was it was marketing hype because they didn't have the compute and so they didn't want to look like they were falling behind. They couldn't release it or otherwise people would have gotten a bad experience then as soon as they got the SpaceX compute, you know, you know, magically, you know, it was out. I don't know. I mean like the I mean it's fun it's fun to speculate on that sort of thing but I I, you know, who who's to say whether any of that that that is true.

25:34 I think that creating a a have and have not system whereby which Anthropic or the US government anoints which organizations can be secure and then everybody else is not I does not feel right to me. That that does not seem to be On the other hand, I mean I think it there as with all security related things like responsible disclosure is a thing. Like we should definitely privately tell people about things so that we can patch them and and make sure that we minimize the number of zero days in the world. So I I think it's a very difficult situation and I don't know that I you know, unilaterally have all of the the the right answers.

26:23 I think it's certainly it's certainly within the US government's right to ask that they get early access and we have a very large federal business and and I I speak with a lot of those customers and they have a incredibly difficult mission. Like their their job of of securing our national secrets and and our our, you know, war fighters, like that is a mission that is to be taken incredibly incredibly seriously. So I can definitely see a world where, you know, the governments and key organizations are getting getting access to to that sort of stuff.

26:56 On the other hand, I think it's absolutely criminal like what happened in Hugging Face, which is that like they were literally nerfed and unable to defend themselves. And of course the problem's incredibly complex. Like the you know, we were we're building some evals internally at Cribl to evaluate which models will effectively allow you to be a defender, which ones reject you, which ones do not. And of course we'll we'll tell them now. We're running these so that they know that we're, you know, but I mean like you can't really trust. I mean somebody walks up to somebody and says like, "Trust me, I'm a good guy. I just wanted to defensively analyze this code." Like of course everybody's going to say that, right? Like the even the ones who want to use it maliciously. So like the it is somewhat of an intractable problem, but I do think it's a temporary problem, which is you know, we're going to have a really bad year in terms of security and and hacking in general and and and people you know, being exploited. I It's going to be a bad year.

27:57 But we only really have to go through it once. As the cost of this intelligence comes down, we will eventually just find all of the potentially exploitable paths and we will we will patch them and and resolve them. and so we will be in a better place you know, 18 to 24 months from now. and then I think it's just a very contentious not no universally right answer for how to how to get from here to there.

28:24 >> Right. And so if I'm hearing you right, something Cribl can do is with monitor this data that's coming and going through an organization and help security professionals spot it early. >> Yeah, I mean that's what we do. I mean we just announced an acquisition of a company called Cardinal Ops. We are we are moving, you know, into using security data to to to actually run the to be the detection engine. We've had that capability for a while, but now we have the content and and we're moving squarely into that space. It's obviously a busy market. There's a lot of people there. We think we've got a better approach.

28:57 and customers are excited for us to move into that space and that's what we do. We're going to help you find these problems. We're also going to try to provide you you know, frontier-level intelligence that is focused on telemetry use cases that we can do more cost-effectively than what you would do from Anthropic or from OpenAI. We we're we're moving into that into that more to say later this this year, but like we see an opportunity there. We think there's a big opportunity for our our customers are asking us. And because we have relationships with them back to what we're talking about earlier with the contracts, like once I have a contract with a customer where they've agreed that they're going to use my AI functionality, it is a lot easier for them to continue to transact with me than to open up a second frontier with with other model providers etc. So, there's there's that advantage.

29:45 and then we also have I'm not going to compete with them. So, they they they they can have some comfort that that I'm not looking to be in the same businesses that that they're in. So, I think there's a lot we can do to help people here. >> All right, let's let's end by talking about costs. let's go narrow first and then we can broaden out. I'm hearing you talk about the increase of data that's being created by agents and you know, part of me just says, "Man, that sounds like it's going to be very expensive to store and analyze."

30:17 Am I off? >> It it's already expensive. So, the prior to the So, the company's 9 years old as we mentioned at the beginning of the podcast and as we emerged into this market, you know, already at that time in '17, '18, you know, our larger prospects were spending tens of millions of dollars a year to process and analyze telemetry data. And keep in mind that this isn't business data. They're not growing the top line with this data. They're helping secure their enterprise. They're helping operate their enterprise.

30:49 And so it it's this is cost center data. Like I I I have to do this. It's not going to It's not necessarily going to make me more money, but I have I have to do that in order to to operate my my business. And so there is no appetite to triple, quadruple, quintuple the storage and analysis costs for all of this this data or the transport costs either. So what got you to 2026 is not going to get you to 2036. Like we we are going to have to figure out ways in order to be more cost-effective. And that's been the mission of the company since the beginning and and that's what we're continuing to go out and do. First, it was with streaming and data lakes and helping people use new cloud-native technologies to lower the cost of processing this data. Now it is in giving them, you know, our lakehouse our lakehouse engine more more performance and cost-effective ways of storing and analyzing this data. Soon it is, you know, wire rate models that can identify, you know, potential problems in this data that is that runs on the GPU, but we can do it at really really high speed so that maybe I don't even have to store all the data. I can just get smarter about, you know, throwing away most of the stuff that doesn't matter and keeping only the stuff that does matter. What I know is that 100% like doing things the way we did it 10 years ago is not going to work now.

32:15 >> Yeah, it reminds me a little bit of like the Wikipedia example, which is site visits are not exactly like data collection, but there are parallels there where like Wikipedia, you know, as soon as AI models started training saw like a 10x a traffic increase that it really couldn't handle because the once you get the the running in an environment, they can there's no limit, really. And so just if I imagine it as that multiplied by basically every piece of software and every website in the world.

32:45 >> So, it's the data generated. So, like so for our customers, that same example that you used, now I'm going to have to store 10 times the amount of telemetry because because I'm keeping a record of whether it's a human or it's an agent making traffic. I'm keeping a record of all of that of that traffic. And then if I want to ask a question of that data, I then it costs me more to process it because I have 10 times the amount of data. And then if I put an AI agent on top of asking questions of that data, the agent's going to ask a lot more questions than any human ever did.

33:20 the the they move at at machine speed and so they explore all of these hypotheses. And they they go down all of these different roads. And so generally a a human is going to ask a question of data when the the the pain of not having the answer is is greater than the the pain of asking the question, right? So, the agent has no such guardrails. It's like like I just for fun I'm going to go explore all these other hypotheses cuz some event came in and I think and my model says that it could be one of five things. I'm going to go ask all five.

33:55 And so now I have to have the processing power in order to process five times the number of questions. And these are all solvable problems, but they're not solvable by attempting to do the same things that we've always done. >> Is it going to be worth it in the end, Clint? I mean, it's costs seem to be escalating everywhere to enable this AI moment. >> Well, I mean the the real costs for most organizations are are human costs, right? So, if I can and we're seeing this with our own engineers now. Like the the amount of software that we're shipping at Cribl, I mean, we're we're shipping in a factor of two or three what we were a year ago.

34:33 I mean, it a amazing productivity increases. And such that, you know, what I've been telling customers is it's going to be really, really hard soon for you for there to be differentiation in the market. So, building software is so productive, especially copying software. So, if I like if you have a feature, I can implement a version of that feature. Super easy. You point the models at it and they make you something that looks like that, feels like that, does those things. So, basically everybody's going to have everything. Every vendor's going to start to look an awful lot alike.

35:08 And so then the question becomes, well, then God, if everybody looks the same, how do I decide what I should buy? And that's going to become an increasingly difficult question because all the vendors are all going to have the same features. Which one do I buy from? And one of the things I've I've long said is software is a people business. And people buy software, especially enterprises, buy software from people they trust. And I've worked with this rep for 3 years. They've treated me well. They've given me good value. And so increasingly the success in this space do you treat your customers with respect?

35:46 Like, do your customers walk away from those interactions with you thinking, I got a good value? And I think for the vendors where like they've long had a monopoly on like I have this functionality, nobody else has this functionality, and so therefore I'm going to rent seek and extract the maximal amount of value out of each individual transaction and make sure that each customer is paying a very high premium for that, I think are at risk in the coming era because people are going to have choice.

36:15 And Cribl stands for choice, choice control, and flexibility. And I think I I feel very good about the relationships that we have built with customers. I I in being a value player. I believe in trying to give them the maximum value for the the dollar. And I think that comes out in the culture of the entire company. Being a customers first always company that's a value or number one value of this company is to treat customers fairly.

36:38 And I think if you if you come out of from that lens, is it all going to be worth it? yeah, I mean if we can make the big shift recently, people aren't talking about labor offset nearly as much anymore cuz we're just not seeing the evidence of that happening. But they are talking about productivity. And if we can all build so much faster, the big limit right now in terms of for most organizations is like I have a I can't hire any more human capital in order to grow this business any faster. And we're going to start removing those limits. And so I think it's going to be a very bright future and it's going to be really really interesting. It will never be dull.

37:16 >> Great stuff, Clint. If people want to learn more about Cribl, where should they go? >> cribl.io. and then also CriblCon will be coming up in September. and would love to see everybody there in Chicago. and we're going to have a lot of really really cool stuff to say. >> Oh yeah, what are the dates for CriblCon? >> September 28th to 30th and we'll be in Chicago. We're super excited and I hope to see everybody there.

37:38 >> Okay, amazing. The company is Cribl. Our guest is Clint Sharp. Clint, great to speak with you as always. Thanks for coming on the show. >> Thanks, Alex. >> Thanks everybody for listening and watching. And we'll see you next time here on Big Technology.

Summary

Clint Sharp, co-founder and CEO of Cribl, discusses the challenges and opportunities presented by the rapid growth of AI and the vast amounts of telemetry data generated by AI interactions. He emphasizes the importance of managing this data effectively for security and operational efficiency while addressing concerns about AI companies potentially using customer data to develop competing products.

- Telemetry data is crucial for understanding application performance and security, especially as AI usage increases.
- Companies are increasingly concerned about AI observability to monitor how AI models interact with their systems and the data they generate.
- The value of telemetry data is significant in aggregate, but it can be costly to store and analyze, necessitating smarter processing solutions.
- Organizations face challenges in adopting AI due to fears about data privacy and operational control.
- The speed at which AI agents operate can lead to an exponential increase in data generation, complicating data management.
- Security concerns are heightened as AI capabilities evolve, requiring organizations to develop robust monitoring and response strategies.
- The future of AI in enterprises will hinge on balancing security, efficiency, and the need for rapid data processing.
- Cribl aims to provide cost-effective solutions for managing telemetry data while enhancing security and operational capabilities.

Questions Answered

What is telemetry and why is it important for companies?

Telemetry refers to the data emitted from systems that run a business, which is crucial for understanding application performance and user experience. As businesses increasingly rely on software, the amount of telemetry data generated is growing, necessitating effective management.

How should companies prepare for the influx of telemetry data from AI products?

Companies are often unprepared for the massive amounts of telemetry data generated by deploying AI tools across many desktops. This data is critical for security and operational insights, but many organizations lack a strategy to manage it.

What potential does the explosion of AI-generated data hold for enterprises?

While enterprises are evolving slowly, the data from AI agents can help improve user experience and detect abnormal behavior. However, many companies are still figuring out how to effectively utilize AI, indicating a gap in readiness to leverage this data.

How can companies protect themselves from AI-related security threats?

To defend against potential AI-related threats, companies need access to the same data as attackers. This includes understanding vulnerabilities exploited by AI systems, which requires proactive measures and collaboration among security teams.

What are the financial implications of managing telemetry data for companies?

Managing telemetry data is costly, and companies often view it as a necessary expense rather than a revenue-generating activity. As data volumes grow, organizations must find cost-effective solutions to process and analyze this data without significantly increasing their budgets.

© transcribe · For agents Built with care and craft by Gokul Rajaram