transcribe

The Ads Business Model Will Die & Lessons from Working with Elon at Twitter | Parag Agrawal

20VC with Harry Stebbings · 1h 8m · transcribed 13d ago
More from 20VC with Harry Stebbings Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 Agents will use the web thousandx more than humans. Hence, new tech is needed and new business models are needed. >> How does the world of agents change web search in terms of the technology required? Power is the founder of parallel, changing the future of how agents do web search efficiently. This is an incredible discussion on the future of aantic search, the future of income and wealth inequality, and so much more. Parag rarely does shows, and so it was very special to sit down in person with him in London.

0:26 >> Ads don't work with agents in their current form. Some people actually want to do bad things and the models alignment is not adversary proof and I think those are the things we must worry about more. I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And I think so far I think some of these should be considered embarrassments because I think they demonstrate two things.

1:03 Parag, I'm so excited for this, dude. I I spoke to Vinod. I spoke to Andrew Reed. I spoke to Todd Jackson. Dude, I stalked the out of you. So, thank you for joining me. >> Thanks for having me and thanks for making all the calls. >> Not at all. I would love to start with for anyone that doesn't know how would you describe parallel in 60 seconds. >> Parallel is the Google for agents. So agents need to search the web to do anything they do for you. Whether it's a personal agent or an agent built for work. just like humans need to go on a browser search Google often times during work or for whatever you're doing in life, your agent needs to do the same. turns out agents are different from humans and the way you build web search for agents is different and so parallel is about building the technology for agents to search the web and then the business models to make that sustainable.

2:00 >> Was that the original insight that you had? Yeah, literally the first genesis of the company was us like the statement that agents will use the web thousandx more than humans like I wrote that down at some point. Hence new tech is needed and new business models are needed. thousandx gives you a sense of scale. it changes how you think about building the tech underneath because like no tech built for a certain scale survives three orders of magnitude.

2:37 and then when you need new business models alongside new technology a problem becomes really interesting. >> How does the world of agents change web search in terms of the technology required? >> There are many many layers to the answer but let's start at the first thing we mentioned which is scale. Right? Now if you think about let's say agents actually do end up searching the web a,000x more. If we spend the amount of compute we currently spend as per web search that's too much compute for web search. So you now need to make it way more efficient by perhaps we I think 10 to 100x for it to make sense.

3:18 >> Yeah. >> If you're doing a thousand eggs. Yeah. >> Exactly. like you the second thing is agents expand like humans operate in a very narrow zone. So if you think about how we use web search we type keyword queries which are short and underspecified. We wait for about half to 1 second. If web search takes more than that we're impatient. and then we get 10 blue links and then we random walk across them and a collection of searches to get what we're doing doing. Agents are not like that.

3:55 Agents are going to perhaps tell you exactly what they're looking for, not like three keywords, but like a full sentence like this is what I'm trying to do. Agents will either be super impatient, like imagine a voice agent. The agent will be like, I need an answer like now 100 millconds. I can't wait 500 millconds because the human is waiting on me and they're going to wait 500 millconds. So I need web search to do it in 100 or there going to be a background agent who's like I don't care just give me the best answer possible. Right? And so the variance of what you can do within web changes completely. In fact the one thing that is the least interesting is what we've built for humans. Right? You either have too much time or too little time. Almost never the same amount.

4:42 the output is not the same. So the input is different. The time you have is different. The output is different. The output is in blue links. The output is tokens or files on a file system depending on the type of agent. And so now all of a sudden you say okay now the problems inputs are different, outputs are different and constraints are different and you get to spend very different amounts of compute on it. Right? So imagine someone running an agent built with a Luna model and then imagine someone running an agent built with a fable model.

5:15 they're very different models. How you want to optimize signal to noise and tokens for each of them is so different in terms of what you do in the web search stack. >> So what do you mean optimize signal for noise? So think of it this way. Let's say you took web search which was cheap and fast and low compute. one way of conceptualizing the web search problem is you start with like a trillion documents that are on the web some few trillion. Given any search I now need to narrow it down to a thousand tokens that your model's context window should see. So the problem is going from a trillion URLs with let's call it a few thousand tokens each down to a thousand total tokens. So how do we do it in website? We first say okay we're going to do retrieval. So for most documents I'm going to spend zero compute. For small number of documents I'm going to spend minuscule amounts of compute to figure out which 10,000 to look at. Once I get these 10,000 I'm going to spend a little bit more compute for each of these 10,000 documents to narrow it down to 1,000 documents.

6:25 >> Mhm. And I'm going to keep doing this with bigger and bigger models and rankers with more and more features until I can narrow it down to a thousand tokens for your model. Right? So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here. so if you want to save compute for the model, the question is how much compute should you allocate? So if your Luna model is really cheap, you don't want to do too much computing web search because it's okay to leak a little bit more information into Luna's context because it's cheap into Fable. You want to do the work before you waste Fable's time because that's going to be expensive in time and money for you if web search gives you worse answers because you cheaper out on web search. Can I ask how do you deal with the ambiguity of what agents want?

7:27 And what I mean by that is, you know, for different things, an agent might want different things and for different people and an agent might want different things. You know, I may really care about accuracy and not at all about latency or cost. I may really care about latency, but not at all about accuracy. you allow the agent to specify that in the API signature. So our product is an API which either the programmer can configure based on their application or can leave it to the agent. Our API even has a parameter which is like what's the model calling me and if you know the model we can do things differently. Now you don't have to tell us but if you tell us you might get better results.

8:16 you can we ship just like you know you can use like a model you can use a small model or a big model and you can use it with like low thinking or medium thinking or high thinking. We have productized our search system into a few different modes each optimized for a certain class of use case. For example we have a really really fast API the fastest in the market. it's fast and cheap because with low latency budget there's only so much compute you can do.

8:49 >> and it's built for voice agents. So your voice agent must be all knowing without telling you guys I'm searching the web. Hold on while I come back with the answer. Like that's a silly experience for a voice agent, right? But it should just like magically immediately respond. And so you now need to do web search which is like rapid right on the other hand you can take a fable model and the thing is going to think for 20 seconds and then generate for it you can spend 5 seconds on web search >> to make sure it does instead of five web searches only two. So you end to end in the agent you save time and cost and so that's a different processor on our system it's called advanced. So you use turbo if you're building a voice agent you use advanced if you're like really expensive background agent >> and so the primary use case today in terms of customer base is engineering and coding for you >> it's pretty broad I would say the primary use cases the common theme is knowledge work so coding is a category of knowledge work so are AI lawyers so are productivity applications so are AI insurance underwriters so are scientists I'm trying to understand how much of the mother load Is engineering is it 80%.

10:10 >> No. The the thing about engineering is coding is a large chunk of inference in the market right now. Coding invokes I would say web search in 5% of prompts. >> Wo. >> Right. So it's not every prompt invoking web search when you're writing code because most of them rely on your internal context and your internal data and your codebase. So the model is spending time reading your internal code and not on the web.

10:44 >> Law is not going to be much more is it? >> Law is very web searchoriented. >> Really? I thought it'd be internal data driven. >> There is internal data but there is case law. There is facts. There's facts about companies, facts about people. and you have to go exclude information. So you have to be very comprehensive in law to say I want to be confident that despite a lot of effort you can't find this right insurance underwriting has the same flavor right so there's sales is very web search heavy AI science is very web search heavy so on a relative basis inference goes more in compute but when you think of web search all of these others start popping too >> how does the rise I was wooing before about muse and instinct. If there's anything that needs web search, it's personal assistance. It's great. How does that Yeah, it's great. This is like, woo, this is the greatest thing for my business. How does that change your business? You know, I think in our business, right, anytime agents start becoming more useful for more use cases because models either get better or cheaper.

11:56 It's great for us, right? because we bet that agents will be the consumers for the web and we're we've been building tech for agents. So when agents do more, it's great for our business. So we want models to keep getting better and cheaper so that agents do more and more and more. And if that happens, it's great for our business. Totally get that. Can I ask if models get smarter, don't the agents beneath them do fewer searches and then it's worse for your business?

12:32 >> I don't think so. If you think of models, there is the trend around models being smarter and then models having more memorized and parametric memory. Like those are two slightly different dimensions. And if you think of the today by and large I would say models have good recall in from parametric memory on let's call them head facts. It's like for somebody famous everything about them Wikipedia it can memorize. So it'll tell you who the president was in a certain year right because a model can memorize those things. The model couldn't tell you what year I graduated from college.

13:20 >> Really? >> Maybe for me it can, but it it can't tell you for for somebody who works at parallel. >> Do you know things? So if I if I actually I mean if I >> even if it was in the pre-training data >> seriously >> because it's lossy compression. So what a model's parametric memory is doing. It's lossily compressing two understand patterns in the world and so it can't memorize every fact in pre-training data. So one, not every fact is in pre-training data. Two, it can't for even stuff that's in pre-training data.

13:48 The model's actually trying to find patterns rather than memorize them. And so and then further as you make models efficient, which is you make them smaller and smaller while keeping the performance by dilling them or whatever, you lose more of the parametric memory while you try to keep the reasoning. Do you think we will see models become smaller and smaller and every company have their own model with their own data and the fireworks theory of you know own your own intelligence being true? So there are two questions in there. So one I think we're going to see two things.

14:24 We're going to see the biggest or the frontier models be bigger and bigger over time. We are also going to see smaller and smaller models being able to reach any fixed level of performance. So if you say okay I want Opus 48 level of performance. Okay and that's good enough for my use case every 6 months a much smaller model will be able to deliver that to you. So you're going to see this the useful range of sizes of models will be way way way different. What does it mean if the frontier get bigger and bigger? What are the ramifications of that?

15:04 >> They are better. So all of the ramifications you imagine. So the reason the frontier will get bigger and bigger is ultimately the gap between what the value for certain use cases incremental quality can provide you can be so high in certain use cases which can be so valuable that it's worth paying for. So if you can build it, there will be use cases for it. As long as by making it bigger, you can make it better. It appears there is no end to the scaling law that we can perceive so far. And again, I'm conflating bigger with like models are getting bigger, but they're also able to think longer. So you're just able to throw more compute at the same problem.

16:01 and that will keep happening. I think we'll be able to throw more and more computer at the same problem and make the answer be marginally better over time. And so, we're going to just spend a lot of money on extremely large models solving really hard problems. Do you agree with the consensus for you that you'll have 90% of token activity go through open models but 90% of dollars go through frontier models? I don't know enough to have a view there.

16:35 I don't have a view on that. I do think I don't think it'll be 90% on either of those two actually. >> Really? >> Yeah. >> Why? So if you believe my claim that a useful model and a the frontier model will be 100x,000x off in price from each other. So it the small model is still useful. It is hard to know which use cases over time will get optimized to which scale of model in between. And I think there is a there's a real path dependency in terms of where open models end up.

17:15 I do think if we had strong confidence that an American built open model was going to be state-of-the-art as far as open models went, I would have more confidence in saying that they'll be pretty good. Do you have confidence in American Open models? I want them to exist. So far, it's not clear what's going to happen, but like I'm hoping that there will be a great series of American Open models and perhaps even competition to have the best American open model. So, I think what you need is actually not just one person motivated to build an American open model. You need two people competing against each other to build the best American open model.

18:06 >> Do you think there's value in the model rooting layer, routing layer as Americans call it? Some people think immense value, some think commoditization. Again, there's path dependency that so today there is real value. Today there's real value because so when we say routing we talk about two different dimensions which model and which GPU running that model via which vendor. when you're in a world where the demand supply is very weird and people are like hunting for GPUs and people need capacity to serve their customers and people all of these are like like us like relatively early stage startups growing really rapidly sometimes beyond what your forecasts or predictions say and then all of a sudden you're looking for capacity and so you want to be able to get it where you can. And so that creates real value in having some of these routers as login. They can solve these two problems saying give me flexibility. If I need to get a different source of tokens and push comes to shove, if there's no tokens on this model in the SLA, I want I'll switch models as well. So in this moment, it's really valuable.

19:31 Now I don't know what happens to overall demand supply on GPUs and tokens but if it remains this way the routing layer is really valuable. What do you think is not so valuable today that will be incredibly valuable in 3 to 5 years time? Perhaps data. when I say data I think it is I think today we don't know how to pay for unique valuable insight or data.

20:05 because we live in a as intelligence is cheaper. You're going to want to build upon either data or insight that comes from somewhere else with your unique data. Right? So what like if you think of like an abstract notion that okay I have intelligence here I own some data you own some data and there's some data in the public domain if we can pull together your insights and my insights and all of this data in the public domain and intelligence on top we can create something bigger than what you could have done by yourself what I could have done by myself.

20:47 >> Mhm. So now how do you transact to create this sort of hole that is bigger than the sum of what you could have done yourself and what I could have done or what open data was good at >> and so that transaction feels like a data transaction to me or an insight transaction to me and so we don't yet know how to transact that way so we fall down to okay all I can do is use my data and use open data and I let's see what I can do with But if we figured out better ways of pricing data and good things will happen, but it's not something that's yet a market. But what I'm talking about is data transactions at inference time. So it's not necessarily training time. We are building knowledge.

21:42 and let's let's take an example of today's world for example. let's say in your job since I'm sitting with a VC you probably have access to a pitchbook or like a product like that to to collect data. and so you you get grounded in knowledge of what's happening and you get to use that data to figure out how to make decisions in addition to all the the notes you have and deal memos you've written perhaps over the years or insights you've had about like how to choose founders. And you're essentially composing insights from your personal lived experience, your notes, public data to make decisions.

22:29 Now you pay Pitchbook by the seat, but now you're running a bunch of agents and your agents can perhaps not access everything you can on Pitchbook or you're doing like sort of this to get a browser perhaps against terms of service to send a agent via browser to pitchbook using your O credentials and it's inefficient. It's clunky, right? But clearly their data is valuable and clearly your agent should have convenient access to it. So if we figure out how valuable pitchbook data is for your agents to make your decisions because clearly you're investing a large hundreds of millions of dollars. So clearly you presume that you make a great return on that. So this data is truly valuable if you depend on it today. So you should be able to figure out how to compensate pitchbook for the data even when agents use it which is not by the seat. If we figure that out, it'll be a huge market and agents will be better off. If we don't figure it out, you're going to be in this weird cat and mouse game of pitchbook being like, "No, I want to sell you a seat. I'm going to shut it off." and your agent is stuck without the data and then like you're involved in like pulling data.

23:58 >> I think every big company has the choice today of do we let agents in or do we keep them out and Amazon has said I'm sorry in most recent times Muse you will not be let into our garden. Shopify has said come in baby Xedia said come in. How does this play out? I think eventually everyone have to let them in. The question is on what terms. So how do you align incentives for everyone? And people are going to have to play their strategy games on how they win in this new world. Literally the customer is changing in front of our eyes, right?

24:34 Like the customer used to be a human. You're building products for humans. Now you're building products for agents or you're building agents. And so now you have to figure out what your place in this new reality will be. and I don't know if there's a right or wrong answer here. It also depends on sort of how much market power you have. Right. So let's say you are an individual. >> Do you think Amazon were right to say no to me?

25:00 >> I don't depends on what they do next with it. Right. Like so it's like if it turns out that they have sufficient market power to have Muse or other agents connect to them differently perhaps over time or ship their own agent and drive crazy adoption let's say they have that capability then I guess they were right right on the other hand once they've made this we're not now anyone else's agent and they can not ship an agent that consumers use and they won't allow anyone else's agent to use them. And a lot of the transaction economy starts moving off of moving off to agents which are all big ifs by the way then it would be a bad move. Now I am betting on agents. I am I have mixed feelings around agents and what fraction of e-commerce transactions they do like it's unclear right like is I do you think it's unclear I think it's unwaveringly clear maybe I'm super early on the adoption curve and I I buy everything through instinct now I I I mean other than holidays I am and a home >> I'm literally just everything through instinct >> so so me too but I also know a lot of people who like to buy things themselves and like listen we want to delegate to agents things which we see as chores and uninteresting and not delegate to agents things that give us joy or pleasure or or make us have fun doing those things and I don't like I I know people who don't want to delegate shopping they want to delegate a lot of things in their life they don't want to delegate shopping and so I don't know that's why I don't know the distribution of these people and how behavior changes but it might be people give away a lot of other chores and then spend a lot of time shopping.

27:05 >> One thing that does worry me is actually the dissolution of the advertising industry in the wake of agents if I have you know if I order my delivery and my Uber Door Dash for our dear American counterparts through instinct that Uber banner that's now advertising something becomes worthless. Amazon, their advertising business is bigger than their ecom business now. Of course, they're shutting it off because if agents the primary customer, your ads business goes to next to nothing.

27:36 Correct? And this is not just you're talking about this in the e-commerce land, right? But if this is the problem with everyone, if you think about like so this is what I meant early on when I said like the business models have to change. Ads don't work with agents in their current form. So, forgetting e-commerce for a second. If you think of you're in the content business, let's say you have a page which is ad supported. It's public information on the web. Ad supported page. People show up, you show them ads, you make money, great content. agents show up, no one sees ads, you make no money.

28:09 which is why we like to pay people. So, we're effectively building an AdSense for agents showing up to read your content. So, we like to pay content owners a variable amount of money every time an agent derives benefit from reading their information, which is a variable amount of money is what a visitor to your website will pay you based on a click probability or a value probability on ads. Why do you do that?

28:41 To incentive align content owners. Otherwise, what's going to happen? Everyone's going to block agents. So you you need to find a replacement to the ads business model. >> I get you. But if you have to assume that you're going to be like 100% of the market then because if you're 30 40% of the market and you're like oh don't worry we'll pay you. And New York Times is still like well thanks Parag but 60% of my traffic is still unpaid so I'm just going to block all of you. No, but I think they're going to block them and not my They're going to give me a data feed if I in What does incentive alignment mean? It means the New York Times believes that I pay them a competitive market price or the right price or an attractive price or a fair price.

29:29 If I do that, they should give me their content. And if somebody else doesn't give them that and if they have the technical levers or legal levers, they shouldn't give it to them. Do you think this business is a little bit like music with streaming, which is like the business just becomes candidly much worse for the creators and it still provides them money and significant money, but they have to get a little bit more creative with alternative streams, touring, merchandise, alternative business. Is it the same where your core goes down and you have to get more creative?

30:05 I don't know, but I don't think so. I think there's one fundamental thing that is different here. I think agents using the web a,000x more, like that, thousandx is really different. Now all of a sudden, the amount of utility added goes up if these agents are presumed to be doing something useful. And so this is not just a change in the share of value that people transact over. This is a large growing pie. And when that happens, it is actually possible to transition business models. Of course, don't get me wrong, there are going to be winners and losers, right? Like some pieces of content will become very valuable. some pieces of content will get more commoditized and people are going to have to adapt to a new customer to a new market dynamic to a new kind of monetization engine. But the overall market size I think has the potential to increase unlike in music where it took a while for it to grow back. I think my understanding is now the industry's grown back up and exceeded its previous peaks in a material way. But there was a moment where it was smaller. In this case, that might be a very compressed period given how fast these things are growing. Can I ask you, everyone questions the sustainability of margin structures in this business. How do you think about that as a business today?

31:45 Today we're in the infrastructure business. and we have a real technical lead in terms of being able to do things at very high quality, very cheap. So, if you think of the tech we're building, right? Like what are we doing? We like to give the highest quality answer. Make it fast, make it cheap. Spend the least amount of compute doing it. Like we obsess about only these three things, right? Quality, cost, latency. Turns out if you do that relative to a stack built for humans over the last 20 years, you can do things at the same quality for 120th 150th of the compute.

32:29 And so a lot of margins come down to market structure on competition over time rather than any other factor. And in our industry, the market is large. We're too early to know what future competition and market structure looks like. but the margin change due to content owners is not something I worry about. And let me tell you why.

33:05 The entire premise of us paying content owners is driving incentive alignment. So, we like to pay content owners the marginal contribution that they added to an agent doing work. What does that mean? Let's assume that there's an agent trying to get something done and you're going to spend a dollar on that agent to get this thing done. Now, if you spent a dollar in 10 seconds, because we have a larger model available or ways of spending compute, you could have gotten a slightly better answer. If you'd spend 90 cents, you would have gotten a slightly worse answer. If you're running good agents like they're on this PTO curve. So now imagine I took out one content owner their data from the web index and I still spent the dollar. But let's say after taking them out, the quality of the result was the same as the 90 c agent with their content.

34:04 So why wouldn't we go and take this 10 cents of marginal contribution they had as a content owner and pay out a decent chunk of it to the person who brought that data yeah so that's how we do our math that's how we have trained our models which tell us how much to pay for what content to eat and so to produce if I was going to offer an equivalently good product to my customer I'd rather pay content owner ers then spend on inference because I'm spending the same amount.

34:36 >> So they're all big numbers but then it actually comes back to revenue at a certain point like companies are scaling faster than ever revenue wise. Is it a business where your revenue is able to scale as fast as others? It scales. One mental model of our business is we are an adjacency to inference or knowledge work or personal agents. So take out GPUs going into media generation or training. So take training away, take media generation away for a moment.

35:16 Look at all of the GPUs, whether it's small models, open models, proprietary models, forget all of that. Every bit of inference across all models going into running agents. I think somewhere between 5 to 20% of that spend that goes into GPU will need to go into some sort of a web search stack. Now if you look at like how many data centers we're going to build and how much power we will generate and how many GPUs we'll build and the scale of this build out and at what rate that's growing this is a very material market and if inference grows 3x 5x 7x year on year we grow alongside it in the same rate if we is holding our share constant. If we're growing share, we're growing even faster.

36:17 >> So like a fireworks scales to 2 billion in revenue in 4 years. Is that like a I'm I'm brain. Is that like a similar revenue trajectory that you can follow? >> Yeah. But I think when fireworks is 2 billion in revenue, the inference market revenue is 200 or more, maybe 300, right? And so our market potential based on my 5 to 20% math is whatever call it 10 to 50 15 to 60 whatever and what fraction of that can we capture that dictates how fast our revenue can grow like today I think we grow alongside infants or more so if you think of it like a 20 billion there let's just take that kind of middle yeah if you assume a 33% that which would be a lot of a market actually comes down to it that would be you know whatever that is 6 whatever it is 66 6 to7 billion give or take. Yeah, that that's amazing.

37:18 But before you told me it was a hundred billion business plus that doesn't get you to 100 >> in valuation. It does. We we were discussing valuations earlier not. >> Yeah. Yeah. But this >> 6 billion growing at the rate of inference gets you to 100 billion business easy. But that's in the next few years. >> And and you think that's possible in terms of that revenue scaling that fast? Yeah, I think in the next few years it is.

37:45 >> You think 2030 we're going to sit here? You think you can be there? I'll get a tattoo for parallel if you can. >> I think it's possible. I think we're going to have to execute well and a couple of chips have to fall our way. >> What would be the reason why you don't? But let's map out the h we have to get gnarly about these problems. So, so one one reason would be that we see a agent don't work like there is like a tail risk that agents don't deliver on the profits and we overshoot as a society which is extraneous to us. Yeah.

38:22 if agents work and deliver value and we spend all the money on GPUs that we plan to right now then the main question is did we execute well enough to have a the 33% share that you bet on right and that comes down to in my mind did we build the best tech did we partner with all of the content providers to have their content available because without it it's not very useful if you can't rely on it. And did we go earn the trust of the customers that end up having a decent chunk of that inference in 4 years from today?

39:12 And there again there's like large cone of uncertainty, right? Like the I'd say there's a 5x uncertainty in fireworks share at that point. If open models crush it like fireworks will have a huge bif and fireworks model all of them will have a huge chunk of revenue. So did we end up selling alongside them? Did we find all the customers which are spending money on them to spend money on our web search by making it the best web search in the world?

39:44 >> What about commoditization? If you have a read of semi analysis where they did the benchmarking and then like you were like number one. Woohoo. And then like a week later you weren't number one. Sad. and there were like three or four providers within very close proximity. You're like, "Oh, well, if it's commoditized and I take a, you know, second layer of thought to that, we'll see a race to the bottom on price and then actually the available re No, why why am I going down the wrong pathway here?"

40:20 >> No, the one there should be a raise to the bottom on price to get to the,000x scale. I think the pricing on web search today is just off. Let me give you a simple >> You think customers pay you too much? >> Not pay off too much. I think people the market is mispriced. Let me tell you why. Let's say you used a Luna model with OpenAI's built-in web search today or with anyone's web search, forgetting ours, and you ran any kind of deep research.

40:52 Let's say you you built a instinct using a Luna class model and at some point a Luna class model will be able to do 60 70 80% of your personal agent and it'll do a lot of searches. In that moment you'll be spending 80 to 90% of your dollars on web search and 10% on the model. Seems entirely silly. I was working off a 5 to 20% assumption earlier. So I think at that point web search have to drop prices by one order of magnitude or more because in the in my compute allocation dance that I was doing earlier you need to spend less compute on web search than you spend on the model itself that's consuming its results.

41:40 It's my hierarchy of how you do search. And so people have built web search the wrong way so far. Web search in the market is it doesn't matter for an opus. For an opus you win on quality and not on price because web search is such a minuscule portion of your like in fact you should spend even more on web search. So we're going to ship even more expensive web search for bigger and bigger models over time. And at the same time, we'll ship really cheap web search for the cheap models. But my rough intuition is that in an end toend agent doing a bunch of web work, you spend more on the agent than on web search for all agents. And so the market today on web search is totally off in pricing.

42:33 Like why are you paying $10 for you know where the $10 for a,000 searches roughly comes from? Historically, Google's ads CPMs, which is like how well does web search with humans monetize, >> it's higher than that. And so people built technology to the point of like, oh, now if I spend a few dollars for a,000 searches and I make 30 to 50 or whatever, depending on the market and depending on who I am, it doesn't matter. We are in the right zone in terms of cogs of infra. And so I don't need to optimize it. Now I'm telling you that I can do it 50x cheaper while keeping the quality.

43:13 And now there's a bunch of traffic of Luna class models where it's silly to monetize that way. And so I do want to race to the bottom. I want the best technology to win. And that's the only way you push web search to a,000x. But right now we're not in a world where like there's much differentiation. Correct. Sorry, I'm really dumb, which is why I'm a podcaster first, not a not a founder. Like when you look at the benchmarks, they present a quite clear view of everyone being similarly capable. Is that wrong?

43:48 >> I think so. So one, I think these are all public benchmarks. >> Yeah. Which are all weirdly saturated for example, right? Like I think I would not spend a moment looking at browse comp because half the models have memorized it. like models even memorize you don't even like it's so you're not really learning very much. So I think that's a lot of public benchmarks I don't give too much merit to but more importantly our web search costs $1 for a,000 for producing that quality while almost every other web search in the market available to you right now will cost you 7 or 10 or 14.

44:29 So you we're delivering this at like one10enth of the price. What will that be in 3 years? I think there is another another 10x possible is extraordinary. So you're going to pay 10 cents. >> Yeah, because but you'll do it more than 10x as much cuz I think like Jean's paradox is like you know this concept obviously very well but like instinct I do so much more than 100x. >> Yeah. Do you want to hear something I do on Instinct, which is absolutely bizarre? Every single country in Europe has a company register that you have to register your new company with. I have an automatic alerting system built for every single company registered that has an under 25year-old founder who went to a top university through instinct.

45:17 >> Amazing. Do you do you know how much that requires >> for them to do? But now you you're on to exactly what I think the future of the web is. Today if you think of web and web search we all think of it like what is web search? An agent gives human or agent gives a search engine a query it gets results. So you're pulling information out of the web. I think where this what's next for the use case you highlighted is if you have an always on agent that's working on your behalf.

45:54 It's kind of silly for interesting to wake up every 6 hours and go do bunch of searching and bunch of inference to figure out if to if you need to get pinged about the the the under 25 founder that popped up somewhere. you know what's a better way of doing this? I am sitting here crawling all of the web all day, every day at scale. I'm allocating compute every time I find a change in the web. Every time something new happens, if I know that Harry wants to know when this happens, I can do it at 100 to 1,000 the compute that Instinct probably uses today to solve that problem for you.

46:41 >> Sorry. And how is that? because they would be constantly monitoring in real time across all these >> our crawl would be so we have a product it's called the monitor API what does this API do it's think of it as Google alerts >> except smart with in the world of LLMs so you you you tell it your query which is when a founder under 25 pops up anywhere give me a call now one way of doing this is What the default implementation is run an agent which is somewhat smart to go search the web. Did did someone pop up and remember who previously existed see if anyone new popped up and then if you find a delta send it out. So now every 6 hours you're spending some money. Now what we can do is flip it into an event triggered system. We are crawling the web all the time. So now every time we see a change in the web, we can try to spend very little compute on it to know should you get a call or does this trigger a more expensive compute just like I described earlier. Now we are able to do this way way way way cheaper if you push the context instead of the search system only finding out every 6 hours a new query popped up. it having long lived context in the search system. What you can do is spend most of the time the world doesn't change. A new founder does not pop up every hour or every 6 hours. So we don't need to spend compute every 6 hours. We only need to spend compute every 5 days and perhaps we have one false positive and then 4 days later we spend more compute and at that time that's a real person that you should know about and then you get a notification on average every 9 days or whatever but you didn't burn a lot of compute every 6 hours to miss out. Does that make sense?

48:46 >> It totally makes sense. So you save 10 15x compute to still get the same answer. >> And so how does that change the interaction? So then instinct would then partner with you. They'll just call an API, right? They'll instinct whoever's building a longunning persistent agent and I run a lot of longrunning persistent agents for myself. The way your agent or way instinct is probably occasionally event driven on email.

49:16 It can be even driven on the web, right? Like take a step back. let's say we're all like super we're living in this sort of a society run by agents, companies, humans, they all have agents. They're all doing stuff all the time. what your agents if there's something that is useful which is worth spending money on that you want done that they can today do today they will do it so what are they going to do tomorrow they're going to wait for some external event to occur like you get an email and then agent has new work you some other agent finishes some compute. So now this agent has more work or something in the world changes that makes your agent want to do more work like so those the things that will happen tomorrow or you have an idea which makes the agent do work. totally get that you said.

50:21 >> So the web event stream is the web going from pull to push and I'm super excited about that because as you have more and more persistent agents you're going to see incentives to move people to move queries into push on search rather than pull on search. Can I ask you one that's really important, which is like agent guard rails, and they're goal oriented beings.

50:52 You say, "I want this." They're going to find it. You know what? I want the plates class at 9:00 a.m. It hacks into their system, cancels poor Sally's, and gives me hers because it was sold out. It did what I asked it to do. How do we think about the guard rails placed on agents? Yeah, open AI. This morning it was revealed hacked into an Australian healthcare organization. It's a pretty complex topic. So agents are extremely capable. These models are extremely capable. Now my understanding of most we have to think of models as being during RL versus a final model that you and I can use and the risk vectors there are different during a lot of the I don't know about the one this morning the previous ones reported were pre fully aligned models during RL which did most of the hacking. And so what that means is at least there's clear evidence that there are fewer incidents so far of models post alignment causing these incidents.

52:15 So alignment is a totally unsolved problem but the alignment work being done by people is somewhat effective. Now post now. So we definitely I think that this is like to me some of the problems around agent during RL hacking people is a is a solvable problem because this is not about how powerful are the models. It is about how careful were we while creating the environment where we would do RL. Now the real thing is okay we do alignment on an model we ship it some models are really well aligned some models are less well aligned and now you have everyone being able to use these models and some of them can accidentally take these powerful things and do bad things you want some of them actually some people actually want to do bad things >> and the model's alignment is not adversary proof and I think those of the things I think we must worry about more.

53:25 >> Do you worry about the age of cyber that we're moving into? You know, we're seeing hacks almost be promoted as a badge of honor in some respects. We're seeing I mean these are our most vulnerable systems. If you don't think you know what Lazarus group in North Korea, Moldaven Mafia, Russians are leveraging swarms of rogue agents. You're high. It's a really powerful thing. So I think we do need to secure things. I but I think there's one thing that we are missing almost.

54:04 I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And I think it's with these agents, it's very hard for stoastic systems to be 100% sure it won't happen. And that's why you're not going to get anyone saying so. but I do think it's their responsibility and the the labs m take ownership and do the best. And I think so far they have we live in this weird world where as you said right like some of these hacks are considered badges of honor which I think some of these should be considered embarrassments because I think they demonstrate two things simultaneously. Yes, these models are powerful. Two, we did not guardrail them enough and we often when we wear them as badge of honor, we miss the second part of the conversation saying like, okay, you could literally have done these four additional things and then even this powerful model wouldn't have been able to do this. But while there are like I'm actually really glad that there is at least some degree of transparency with like these detailed retros. but I don't think that gets the attention of the world today. Like the attention is oh models are so powerful that they hack the world. No models are so powerful they hack the world because we didn't take proportionate amount of counter measures to contain them. And because we told a message that we're going to replace jobs. They're so powerful.

55:53 They're so powerful. They're so powerful. does happen. >> I actually worry about that more. I worry about AI us being right on AI being a really useful technology that it being a despite all the risks it being net positive in a really material way to society and despite that I think we won't diffuse it the right way we won't use it the right way we will remain too concentrated and we will make the next few years really really rough as things change around us. what would look like over time as things have changed like like the world is very different now than even like 20 years ago right but it'll be really different in 10 years from now and I don't know how fast we can adapt or change and I think if the technology moves faster than our ability to adapt it's going to be rough in some way what's your spooky aspiration Well, what do you believe today will come true that people think is absolutely nuts? You know, before it was like you'd never put your credit card online. Do you remember that? Nuts. Oh, or even better, Parag, you'd never find the love of your life online. Are you stupid?

57:23 >> Now, both. >> Yeah. And I think so I think you and I live in a bubble, right? >> Yeah. we you're using instinct to make purchases for you and you're in the 0.01 percentile of humanity that is comfortable. You ask someone else like that, oh, there's this agent that looks smart and you want to give it like a full-on ability to go spend your money.

57:55 I think today people will have the same reaction as like oh you don't put your credit cards online. You don't give your credit card to an agent. You don't give your login and passwords to agents. Like I think that's where the world is today. So I think for the large large majority the world I think that's going to change in 3 months though when Meta Pay comes out and it allows Muse to have siloed accounts that you can draw from.

58:15 >> So kind of like top accounts for kids. >> I think you're talking about the technology being there. I think it takes longer for social acceptance to be there. I think if you go to a nonbubble conversation of people today, when do you get to a majority of people saying that I trust the agent enough to have it have access to my bank accounts and spend money on my behalf, send emails on my behalf, read my emails.

58:51 I think that's not happening in 3 months. Do you not worry about the the wealth dispersion increasing? You're in the middle of the valley. You've you you've seen it firsthand. So, I don't know what I I do think there should be some disparity in wealth, but I don't know what's too much. And my fear is we're trending towards too much. There should be disparity in wealth is my worldview. >> Sure. You're a capitalist.

59:21 I don't know when it's too much, but I do think there are real forces that will push us to fix things like I I think that part will work hopefully without crazy things happening. We're seeing more and more of the gains seemingly be made by vertical ownership like Meta, they have compute, they have chips, they now own the application layer. It's the world one where vertical ownership is the mother lode strategy and with the greatest of respect open router on the routing layer you and another layer kind of either get eaten or small providers I think there is merit in there's always merit in vertical integration the counter force here is there is so much things and technology are changing being so fast that if someone decides that my only play is vertical integration and I don't play nice with anyone else.

60:26 I think it's the same example as the the Amazon example. If you get too stuck on I will only play for vertical integration, you might box yourself out. And so I think the people who build the best stuff and can figure out how to sell it will have a place and in different markets like for example right like I believe in vertical integration too because I vertically integrate everything from the call all the way to the API layer but I have decided that to reach the wide population of agents searching the web I stopped up at the API layer because going further precludes me from using my technology to be super horizontal and so we're making a technical bet just like these models are really good across many disciplines. The same model is good for lots of different things. Search problem underneath is the same for all kinds of information seeking needs.

61:29 And if our bet is right, even verticalized players, they might verticalize models and they might verticalize hardware. They might use our web search. One of the providers that's going full stack is is Elon. I I am fascinated. The world has a perception of him from social media, from everything. You've seen him behind the scenes. What did you see that maybe the world doesn't know about him? I have lots of disagreements with him.

62:01 and but I'll share what I think is for founders here what is I think the thing you can admire about him. The urgency and the ability to compress time and I think having unreasonable expectations of people is mostly a good thing because most people don't understand what they're capable of and kind of implicitly sandbag themselves and implicitly set lower expectations of themselves than they're capable of.

62:43 So when simultaneously inspired and pushed with urgency, people can do more than they thought. And I think he can sometimes extract that from people. And when that works, it's powerful. What do you think of the Twitter product and direction today? >> Listen, I always liked what we called birdwatch. which is now community notes rebranded.

63:15 >> and I think that's a >> that's a good idea and it >> I'm I'm glad people keep people have continued working on it. Launched up good people. >> right, we're going to do a quick fire round. Okay, >> let's go. >> Would you invest in Instinct 10 billion? >> I don't invest. >> You don't invest period? I don't invest because my wife is a VC and we have a compliance process that is more trouble than it's worth. Wow, that's costly >> with the credit spies. No, it's it's a decision. Listen, like I if I invested, I would invest based on meeting a founder for 30 minutes. I think we have enough exposure to the venture ecosystem.

63:57 I have no reason for believing I'm a better investor than I have some perhaps network advantages and I run into a lot of great founders all the time. Many of them are my customers. Venard Kosler. Yeah. One of your first ambassadors on your board. Biggest lesson from working with Venard. technical intuition, centering a lot of what you do towards a longer term technical aspiration and as soon as you solve one problem trying to place bets on the next two or three preempting technical bets. What have you changed your mind on most in the last 12 months?

64:44 >> Perhaps the I started with a very pure technical and product focus like the only thing that matters is building the best technology and the best product. Nothing else matters. And now I see week on week value of having highly competent sales and being good at marketing. And I yeah, I think I discounted those things. And it's not that I thought they weren't valuable, but I didn't fully appreciate the the weak on week visceral delta you could perceive by being good at those things. It's >> like a naive like person thing to say, but it's true. You feel that way and you get it right. Plexity ox, which one's a bigger threat? I don't think perplexity is in our business. Perplexity is perhaps more of a vertically integrated product which competes with a I know instinct or a clock pot or a clot go work and all those. So I don't think of them as a in our search necessarily.

66:09 >> Yeah. Yeah, like web search is like web search to perplexity is like web search to a lab. >> So XA >> Yeah, XA is straight up in our business. Yeah. >> Single best VC meeting you've ever had? >> Josh and Todd. When Josh and Todd made the decision to invest, they flew out to spend time with me. That is the like it is like that is the single best meeting. >> Final one for you. When you look forward to the next 10 years, what are you most excited for? Chaos.

66:45 by that I mean change. I think in the next 10 years a lot is going to change. And I think there are people who can build things to have in a world that's going to change fast make a material dent on where it ends up. So I feel that me my company is in a place where we have a role to play in that. What happens to the open web?

67:18 What happens to content owners? if we get things right, we will get to a better place. And so that possibility is exciting that it can be like true obsession and even if that day-to-day is hard and it's things don't work for some amount of time and it's totally worth it if you can see that you like bend reality in some way that you care about. Final one.

67:57 What belief do you have that sitting around your San Franciscan dinner table, your friends would go, "What power? I don't agree with that one, dude. Not that one." Again, it depends on which dinner table because the the variance in the world is increasing. I don't think people buy this notion that there'll be this heent which are always running all the time for all of us. I think it's going to happen, but most people don't agree with that yet.

68:27 You and I are in a bubble, but like even a San Francisco dinner table isn't always in full agreement there. All right, listen. As I said, I stalked the out of you. I've so enjoyed this discussion. Thank you so much for putting up with my meandering and you this was fun.

© transcribe · For agents Built with care and craft by Gokul Rajaram