transcribe

Why We Made Jev — Diogo Almeida, TypeSafe Co-founder & CEO

Latent Space · 2h 22m · transcribed 1d ago
More from Latent Space Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

The Paradox of AI Intelligence

Why can AI solve complex problems but struggle with basic tasks?

The discussion highlights the paradox of AI's capabilities, where it can tackle complex mathematical problems yet fails to automate simpler, economically valuable tasks. This discrepancy raises questions about the current state of AI technology and its integration into practical applications.

  • AI excels at complex problem-solving but struggles with basic automation.
  • There is a disconnect between AI capabilities and practical economic applications.
  • The evolution of AI technology is reshaping technological history.
# 15:49

Ethics and API Usage

What are the ethical considerations for AI API usage?

The conversation addresses the ethical implications of AI usage, particularly in contexts that could lead to harm, such as warfare. While there is a desire to prevent misuse, the speakers argue against imposing restrictions at the technological level, as it could compromise the AI's intelligence.

  • Ethical concerns arise regarding the potential misuse of AI technologies.
  • Restricting AI capabilities could fracture its intelligence.
  • The foundation of general-purpose technology should not be limited by specific ethical concerns.
# 31:38

Future of AI and Technological Change

What does the future hold for AI and its impact on technology?

The speakers express optimism about the future of AI, emphasizing the need for continued exploration and innovation beyond current models. They suggest that the advancements in AI will lead to significant changes in technology and the economy.

  • The future of AI is promising, with potential for significant technological advancements.
  • Exploration of diverse AI models is crucial for innovation.
  • The landscape of technology is evolving, and new possibilities are emerging.
# 47:27

Challenges in AI Development

What challenges does the AI industry face in development?

The discussion reveals concerns about GPU constraints and the need for efficient AI models. The speakers highlight the importance of making AI accessible for experimentation and innovation, likening it to a gold rush in technology.

  • GPU constraints pose significant challenges for AI development.
  • Accessibility and experimentation are vital for innovation in AI.
  • The AI industry is likened to a gold rush, emphasizing the need for rapid development.
# 63:16

Improving AI Codebases

How can AI codebases be improved?

The speakers discuss the potential for AI codebases to become more efficient by breaking down problems into simpler components. This approach allows for better evaluation and verification of AI outputs.

  • Decomposing problems into simpler decisions enhances AI codebases.
  • Improved verifiability leads to better engineering practices.
  • AI development should focus on measurable and evaluable outputs.
# 79:05

Economic Revolution through AI

How is AI contributing to economic change?

The conversation touches on the economic implications of AI, suggesting that while it may not lead to mass unemployment, it will cause significant shifts in the economy. The speakers advocate for AI to enhance everyday life rather than dominate it.

  • AI is expected to drive economic shifts without causing mass unemployment.
  • The goal is for AI to improve daily life rather than overshadow it.
  • AI's role in the economy should focus on enhancing human experiences.
# 94:55

AI Use Cases and Reliability

What are the practical use cases for AI?

The speakers discuss the importance of reliable AI applications and the excitement surrounding new use cases. They emphasize the need for thorough testing and understanding of AI capabilities before widespread adoption.

  • Reliable AI applications are crucial for user trust and adoption.
  • Exploring new use cases can lead to innovative solutions.
  • Thorough testing is essential to identify and solve weaknesses in AI.
# 110:44

Philosophy of AI Development

What is the philosophy behind AI model development?

The speakers share their philosophy regarding AI model training, advocating for a pragmatic approach that prioritizes problem-solving over traditional pre-training methods. They discuss the balance between developing comprehensive models and specialized ones.

  • A pragmatic approach to AI model development is favored over traditional methods.
  • Balancing comprehensive and specialized models is essential for effective AI.
  • Empirical questions about AI capabilities should guide development strategies.
# 126:33

Political Implications of AI

What are the political considerations surrounding AI?

The discussion highlights the political dimensions of AI, particularly in relation to governance and regulation. The speakers express concerns about the influence of political agendas on AI development and its societal impact.

  • Political considerations are increasingly relevant in AI development.
  • Governance and regulation will shape the future of AI technologies.
  • Transparency and honesty in AI discussions are crucial for public trust.

Transcript

0:00 Like how can AI be so unbelievably smart? How can we like solve millennium prize problems in math but still not automate even the most basics of works like like really basic wrote stuff but we have this like supercharged engine of automation that just does not have like the right plugs and stuff to plug into all of this economically valuable work and you know like if the whole company of type safe disappears like maybe it'll take like a year or two for people to like truly catch up. I actually don't know how long it'll take. If if model quality matters, then we are going to be in a very good position for a long time, but it it like it it's done, right?

0:37 Like there like this has changed the path of like technological history. Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you.

1:09 The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that work so hard to bring in space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it. Okay, we're in the studio, a special occasion because this week, Diego, my my good buddy, launched Jev and it's been taking over the complete timeline. how do you feel? What what's it like to be you right now?

1:45 >> Emotionally, never been worse. like a I'm a ragged corpse of a person right now because there's so much going on and I'm like a technical CEO so I have like a lot of fires to fight but like mentally it's like I feel I I say this all the time and I've been saying this kind of for years in my overunder events like I feel like the entire AI field is like one of those like carnival house of mirrors and everyone is just insane. pain and saying the weirdest stuff that doesn't make sense. And it feels like for just this week like I I'm on a better in sync with reality and like oh people see it now.

2:30 AI can be so much more than what was once thought and like yes we are going to make like like an AI based economic revolution is back on the table and this is awesome you know this is awesome. I'm so jazzed the developers get it. It's It's Yeah. And I want to show my internal gratitude to developers and >> I'm so jammed about the community and everything. It's so great. >> Yeah. You were saying yesterday that you decided to prioritize the down town hall and not a bunch of like you know VIP investor type people because you wanted to make sure that they are the people that you get get your most attention right. The engineers, the developers.

3:12 >> Yeah. It felt a little like, oh man, I'm talking to like really important people right now. I probably shouldn't reveal who, but it feels a little bit dirty for me to I I'm like perhaps overly genuine in things. Like it feels like dirty if like in my gigantic calendar events of things to people to talk to. You know, the community isn't one of those, you know, and actually in my ideal world, it would be like community all the time. I was thinking, should I host a town hall while walking to your studio? And I'm like, nah, that's too crazy. Sure.

3:41 >> Yeah. >> Well, you you guys have been hosting town halls on the Discord. Discord is now 100,000 people. your Twitter follow these stats. So, holy >> Your Twitter's blown up. you know, it was it was really funny cuz like at AIE, you were like, "Follow me, please." And then you didn't like provide even your your handle. >> I'm a noob. I'm a noob. >> But no, but like that's like positive aura that like you don't know how to promote yourself.

4:04 >> Someone like called me out when I posted like, "Holy we're all three trending trending topics." And then they're like, "That's a personal feed." Of course. Of course. To you. Yes. Because it's what you clicked on. so, okay, let's Yeah. So, congrats on everything. We'll talk about more details as you have them, but let's for people who are like living it under a rock or just just once like the definitive thing. What is Jev? >> Let me think of a That's a hard one.

4:35 >> okay. And I'm happy to like reask if you want to. I'm happy to I'm happy to like just jam on it. I will say like the first thing that I'm relieved about with this question is now I don't have to answer that question to my parents anymore cuz Chad GPT can just explain it. >> so the way I see it is we need a new class of models. we're not attached to naming that class of models. are the most accurate name we've come up with is system one models. There will be reasons but it's there's a reason why we don't call them decision models because like they will be like system one is beyond that. That's all I can say. we didn't expect this to be our big launch.

5:15 So we have stuff in the tank. >> you should have said low-key research preview. >> it's kind of was right. It kind of was. But we so there's a class of models that we describe them as like machine native system one large programmable. I think these are is the class of models where the goal is for code to be the consumer. So as opposed to pre-trained large language models which are meant for like autocomplete of the internet or RHF models like chatbot instruction following models which are meant to like reply to text or RLVR it's in a weird gray area with RHF like these are meant to have things that directly are consumed by code hence the name type safe. So the thing we really really really want is to have like AI like be as powerful as possible and we think the way to do that is to integrate it with software and we are designing everything you know beyond just the outside the deep internals of the model to be optimized for software. So number one Jev is our first large programmable model or system one model whatever you want to call it. Jev is meant to be optimized for intelligence per dollar. hence the name Jev, you know, Jevad.

6:32 >> Jeans Jevans paradox. Yeah. and it's optimized for intelligence per dollar. I love this debate with people about what is the most important between reliability, cost, calibration, and speed. and Jev is meant to be Jev will be the name of models that will be on the frontier of intelligence per dollar. There's other ways to optimize it like ML or at least if you're good at ML it's all about tradeoffs and you know we are just going all out on that.

7:00 >> Yeah and and to me like calibration is one of the new things that people weren't talking about as much. We've done an episode in the past with Clementine Fia of Hugging Face where they were like, "Yeah, actually they're just, you know, and this is your whole argument about RHF. Is their mode collapsing towards what you want to hear the most or what is most likely instead of like their own internal confidence about a thing? >> Can I soap box on that for a second or cool?" Like I've been heard that your audience is the most technical. So I actually want to get into that. and if I went through extreme precision to make sure everything in our launch video is accurate and real, apparently that's very unusual. One of the things that no one paid attention to was the downsides of RHF, in particular, mode dropping.

7:46 >> mode dropping collapse. >> it's the same thing. It's the same thing. And I I I want to have a blog on this eventually, but I like want to tell as many people this as possible because I think it's a very interesting thing. So the spicy take. I believe in Yan Lun a lot. I think Yanukun's takes are actually among the closest to >> What about this? >> Well, should I address this now or should you go and no later go I actually think that among takes Yan Lun is among the most accurate. but he has this very famous/infamous slide about LM are doomed. Okay, >> you know, like that one where where he like has like a pie chart with like a tiny part tiny little thing and says that as you increase sequence length the probability of it making an error goes in. Yes, this one. This one I love this one because it's one of these things that seems mathematically obvious but is obviously wrong, right? Like it's mathematically obvious but it doesn't empirically hold. And this is my favorite thing to teach people about like where >> what's the disconnect >> exactly? And may I or you want to tell me >> about mode collapse?

8:57 >> Oh no no no. So mode collapse is related to this. >> Yeah. >> The disconnect happens because if you are in a mode covering or a calibrated distribution you are like not you are not overly punished about having outliers. You'd expect like you know something some amount of the time you'd be out of distribution some amount of time you'd be in distribution. That's what happens when you cover the distribution. This was like models before GANs. they made blurry images, right? instead GANs mode drop. They like drop the minority classes and just do the really common ones. And this is why this effect doesn't happen, right?

9:32 Like instead of in order to generate really long strings without making errors, they need to like be extremely conservative because it's really easy to see when an error happens. It's very hard to see when like a subtle thing that looks correct happens. And that calibration is like total poison into like the probability distributions of strings. >> Yeah. >> And it's it's a nuance take and like I I think that this is why this doesn't happen and this is why strings are so bad at decision-m or you know overloading the string models are for decision-m is like a bad time. And while we're on the topic of Han, do you agree that his fix is with which is like a world model like a Japa type embedding thing is the right solve? So bas basically like the one of the reasons that it could fail is because you're trying to reason over token outputs and then and then just looping back again and going keep continuing going until you reach like a end of sentence.

10:27 >> like is that and his solve is jea right which is like joint ambition joint embedding prediction. so like is that the solve or you know like do you have a do you have a take on on that? >> Oh man, I probably shouldn't talk too much about the insides of ML. But I will say that my my brand other than unhinged is practical, you know, like even my take here is practical and like am I a scaling law fan?

10:55 >> depends. It dep you know it's it's like it's a scaling laws tell you how much better you get at a thing for amount in. a scaling law does mean exponentially more resources for normally sublinear gains which looks to be a bad investment unless those like linear gains are like really really valuable. but it's all to me it's all about like what can we do with what we have to make the biggest possible different I can curse.

11:20 >> Yeah. >> Yeah. Yeah. We're we're we're approved for adults. >> Hell hell yeah. >> And also we have a scaling law thing if you want to go into that later. Oh, I I could if we That part is not super relevant right now. I actually if you want to go into my bitterest lesson, I think that that's more relevant. But like to me, I'm all about like pragmatics and I think that the JEPA stuff is really cool early research. I really love awesome research. Is it practical yet?

11:50 probably shouldn't say. but like there's just a lot of I just think there's like so many diamonds in the rough let all over the research world right now that haven't been polished because people don't know how to like do the right task and I think that what our launch did it does it kickstart us as a company like yes will it be great for us as a company yes I think it's going to be like even greater for this direction of like programmatic AI you there there was going to be like a gold rush on top of us for because like software is super charged but I think there's going to be a gold rush parallel to us as well on like all the different ways we can expose things to make software more powerful so people can make even cooler stuff and then we are back to like early internet energy you know and I think that's why like you know the Twitter is just like Jeff Jeff you know it's it's like >> it's inspiring because it's it's like so different than what we're used to which is I'm sorry you can't do this, but we do scaling laws and only the big labs can do it. Right.

12:57 >> That that actually if I I'm going to tangent if that's okay. I think you might enjoy this. >> We're like five five tangents in this. >> Yeah, I get lost on all my tangents. >> This this is going to be horrible for the listeners to figure it out, but they they're going to figure it out. >> Yeah, we could edit this. So, popular thing on Discord that people keep asking me, I haven't had the time to explain it yet, is why am I opposed to safety alignment and why do we not refuse? I'm not opposed to safety as a principle, but I think that safety alignment is generally misaligned with users. and refusal is just like obviously a type error. Like if you're a human being and you're chatting with like you know a bottle or whatever you're cloud coding and a refusal happens like I'm sorry I can't read DNA.py that's an annoying time.

13:45 It's annoying right but you can work with it right and you're forced to work with it because of Stockholm syndrome. I have stories about that too. I I need another tangent deep in here. But like if you ever want this in a dependency running in the background what happens if that refuses? What if someone else is using that dependency? They don't know what that system is. Like like you want the software to just stochastically break because a user sent like a weird message in there. Like that is like straight up insanity. It it's it's coming from a place of like people who do not understand software, do not understand programming and like they are obsessed with like I believe this horseless carriage of like AI co-worker instead of unearthing like the full power of AI.

14:30 >> Fair enough. You want something that is the core kernel that is usable everywhere. >> Yes. Exactly. Like the cognitive core, right? And you need this thing to be like so general, so optimized for its use cases. You want it to be like, you know, you want it to work on all the future use cases, all the weird that people are doing. You know, we obviously didn't train on any of that stuff. Is it surprising that it works?

14:53 No, cuz we trained on weirder stuff, my friend. so but one tangent up about like safety alignment safety alignment makes sense for a product in my opinion for like chat GBT and claude like it what what safety what makes safety and capability alignment different is capability alignment is like about doing what the user wants that is sick for software engineers they want their thing to do the thing and the more predictable it is the less they have to test and play around with it is not anywhere close to that yet it could be but like there's so many more nines of reliability that we want in order to make it so good like a database query that you don't even have to think about it. It is just there when you need intelligence. but safety alignment is like the opposite of instruction following. It's when you want to follow someone else's instructions like openact.

15:43 Exactly. And this makes a lot of sense for a product again like chat GPT should do you you like if they don't want to like do like some not safe for work roleplay with chat GPT that's on them because like maybe that's you know what their users who have like parents and kids want like that's fine but in an API that's nuts, right? like that's completely unacceptable because like people need to like program around this and that is that's so anti-user that it's it it I'm I I I I can be an angry person. So, I should try to calm down.

16:21 >> It's people get your passion and I think it's really good. the the one push back I'll give you is like what if we use it to kill people, right? Like that that is the actual like the not safe for work thing. It's private, personal, whatever. But like yes, like we will use it in war and like that that is something that companies can reasonably prefer their APIs not be used for. >> I get that. I I I think that there's like pragmatic places where that opinion can be held. I don't think the foundation of like a general purpose technology is that place. personally like would I prefer that our stuff is not used to kill people? Obviously.

17:01 Would I prefer it's used for like all sorts of like great stuff in the world? Obviously. Will I put my thumb in the scale for that? Yes. But will I do it at the technological layer? Absolutely not. Because that will fracture the intelligence. Every single time you need it to overfit to some weird stuff, you're fracturing its intelligence more and more. And like these things are fractured to the like they're so darn fractured right now. So and as a furthermore thing to me it's like I think intelligence will be more like a database than a coworker. Like I don't think it's up to databases to add checks on whether or not they're used for like what's something that's not great.

17:41 You know like CIA actually I don't know what the CIA does really. You can imagine you can imagine killing people who are not even bad or whatever. >> Mhm. And like I don't think it's the database's responsibility for that. And furthermore, like a thing that has been weird to me is when people like sign up for our thing on Slack and they're like, "Hey, we're going to deploy this. Can we deploy this thing?" I am just like my brother, we are an API, you are a developer. It's none of my business, right? Like like you shouldn't know what the whole task even is because it should be decomposed into small things. We shouldn't be able to know what the downstream users are doing and that is like a good boundary to give software engineers maximum power. Ideally, they use it for the good stuff and ideally we can like help them and like we've talked about like doing open source and charity and all of that. We have absolutely no time for anything else right now. But like they will get any of that bias out of the technological layer as long as I'm in charge.

18:40 >> Yeah, that's great. while we're on the topic, let's also briefly talk about your privacy stuff, terms of terms of use, which got a little bit of misunderstanding. I just want to clarify that up front. I think it's probably takes two sentences from you about like you will not you're not being that restrictive about your API. Like clearly ideologically you take your role as a platform very seriously. >> Yes. yes. I don't know what you're referring to, but like this was I've seen a couple of things about like benchmarking. Like obviously we're not stopping people from Oh man, I should be careful about what I say. I'm realizing >> No, you said you said it publicly that that was in the preview period. You didn't take it out for the launch and now you're going to take it out.

19:17 >> The team is doing stuff that I'm not even aware of. So it's great to know the team communicated that. I asked them to check in with the lawyers about that. Like we are obviously not stopping people from doing that type of thing. I'm extremely in favor. So I'm extremely anti-public benchmarks. I'm extremely in favor I'm medium about private benchmarks that are proxies. you know >> so are you worry about saturation or training on public benchmarks?

19:42 So it's like easy to cheat. >> not only is it easy to cheat there's a lot of in so I think that we are or or anyone who's like competition with us that you know vaguely there is like you could say like >> we got 50 Jeff clones. Yeah. >> Well, sure. So the let's say that there is competition or let's just say that there's let's just assume that there's an industry 2 years from now of people who are doing similar things to us. The thing that we are selling is intelligence per something per like dollar or per second. the no like people obsess about the cost and the speed. I believe that that is it's cool but like the thing that matters is the intelligence. Like the cost and the speed are like are bad things. you know, you're paying them for something and you need the thing back and the intelligence is what truly matters. The problem with intelligence is that there's a geniqua to it, right? Like like the the good model smell like the thing that happened after we launched of like 2 hours later that actually went way bigger than the video which was like holy >> this is actually usable. Well, it's you know, you know, like beyond that, you know, like the >> the launch was crazy and people could really sense how hard we care about that and that's truly what I think the long term of this is and I think public benchmarks are antithetical to this like they are a way to get people trust in intelligence because intelligence has a genic but the public benchmarks are extremely extremely gameable. Even if they try not to, they still will. You know, like back in the old days, every lab had a team to collect data that looks like MMLU to make it look better, which is just benchmarking, benchmarking with extra steps. So I I believe that in the long run it needs to be vibes and trust until you put it into a workflow and evaluate it for that workflow and measure it and have your own sense of like how it does on the exact workflow that matters. And our job is to keep moving the nines of reliability. This is like an everpresent part of our of what we need to be doing as a company and we need to do everything to have people know that this is something we care so much about. You know, like if we wanted to, we could have released Jev like a year and a half ago if we wanted to be dumb.

22:03 >> Oh, like like the you know my bitterest lesson, right? Like architecture and Yeah, >> I'll bring it up since you since you talked about it here. >> Hell yeah. like you know Sutton says that algorithms beats compute very roughly. data matters way more than compute obviously and doing the right task having a northstar is is is the hardest most important thing. This has happened in LLM land twice so far right maybe 2.2 times. You know there's RHF which like shifted the task to instruction following. No one realized that that was possible. RLVR did like a tiny little like edit to the to the direction and now us right RLCD we have a new task and the goal is you know programs in the loop and yeah data matters so so so unbelievably much like I can't I can't emphasize it less >> yeah you consider yourself a data lab rather than like a model lab is that something that's the wording that you guys use >> we are we will always like care so much about data. to me model capabilities means data. Data is so unbelievably complicated and that is what gets nines.

23:19 Like you have no idea how how much data can shift everything. Data is so important. Yeah. >> Holy crap. >> So if people are looking for a job we are hiring infinite data people actually infinite. >> What is a good data person like? you know clearly somebody who cares about reading through the the transcripts of whatever you said for example that you all your data is synthetic y but that's only like that's scratching the surface right like it's not like synthetic so what right synthetic but we have a people with a lot of taste and a lot of care looking at looking at these articulating what's wrong going back regenerating is is that what a good data person is these days >> let me try to figure figure out how to like it's it's super complicated and like I literally onboard the data people with a talk that I assume is longer than this podcast will end up being. So I will try to say like the high level of it. So number one we don't do the kind of synthetic data that people kind well I'll do actually number is zero. data and synthetic data depends on your task like the shape of your data the shape of your task changes the data like RLVR's data is kind of environments right yes RHFS is the human feedback you know each task has its own unique kind of data and we of course have our own unique kind of data right so number one we have that number two the thing I the reason why we don't want to train on our users data even if we could right? Like we could probably ask for any terms right now and it will we I don't know if it would make a difference. We truly don't want that.

25:00 because no matter what the real world data has so much bias. There's like a power law of like people like asking the same things where you'll end up like overfitting to it and like fracturing to it and all of that. And number two, we are like aiming for like a complete sci-fi future years from now where like these models are going to be like the general infrastructure layers and layers and layers and deep down the stack to to like things people can't even imagine.

25:25 Like I would like to think of our model like kind of like like you know UDP as LLMs and TCP as our models. All sorts of stuff can be built on top of that. And we need to be able to nail those futuristic use cases such that software developers can actually build that futuristic stuff. And the way to do that is even if we had all of the data of the present, we would just overfit to the present and then it wouldn't work. What we need is to like it almost feels like like they're like artists, you know, they study this cognitive core. Our cognitive core is like way less jagged than anyone else's. And then they find the jaggednesses and then they address them surgically in a way that and you can never perfectly do this right but they do it in such a way that it addresses it in every single possible like dimension >> general case rather than the >> exact and like that requires a lot of intelligence every time.

26:19 >> okay so we mentioned a little bit you you sort of criticize my thinking as like very RLVR influence which is like very fair. let's actually mention RLCD. which obviously you have some secret sources. To my knowledge, you've never actually published a paper or anything like that on it. No. Right. >> No, not yet. >> but like what what should people get from this? Like what can you give people some confidence that you're just not just making up jargon for the sake of sounding cool, right? one thing for me is like calibration I do think is it to me like well understood because we've covered it in on the podcast but I I don't know what you mean when you say RLCD versus what people are >> familiar with it's a great question and actually I will give a related question what is RHF >> right and actually RHF means multiple different things right like there's the RHF of the original I think it was like Paul Cristiano teaching a robot to backflip or something like that wasn't there something >> was that it >> that was the original the PO paper, but I don't I don't know.

27:15 >> and so PO is not necessarily from human feedback if I recall. Okay. but I I believe it was like an open alignment work that could teach hard to specify outputs like a backflip. I'm not 100% sure. and then there was actually learning to summarize. You know, this was work by a bunch of the the team that helped with instruct and co-authored the instruction following paper which was teaching doing PO on language models. >> This is the sorry I'm trying to trying >> trying to manipulate this this thing.

27:50 this is 2017. >> Yeah. I'm not 100% sure, but like that looks quite right. If it has like a robot doing back flips or something like that, that might be it. >> Yes. Okay, cool. I guess I I got it right. Hell yeah. >> there you go. Yeah, that's the one. >> So the idea was can like can you do like illsp specified things with it? So that's like version one. Version two was the learning to summarize work that like open did which is actually like po on language models to do something somewhat illspecified. This is like another thing that people refer to as RHF which I did not co-author. Oh, Dario's there. Cool.

28:29 >> hell yeah. >> And Rafford. >> Yeah. Yeah. Yeah, shout outs to Alec and Ryan. Love them. but the thing that I refer to RLHF is the Oh man, >> of I have comments on that. >> I have comments that paper, but like we're so at tangents deep. Yeah. so the thing that really got to to me, the thing that I'm calling to RHF is the task of instruction following. It's not about the PO, that part doesn't matter.

28:56 It's about like setting a north star of this is a valuable direction. It's kind of like the bitterest lesson lesson northstar and for us RLCD is this new task. >> and it is not I I don't see it as jargon like I I try to communicate with precision. It's just that hey here's another northstar just like DPO and all of its like you know descendants also do RHF despite not using the algorithm in that paper. Mhm.

29:24 And so clear clearly stating the northstar is being pro programmable AI is is is one word that I really catch on to removing the human in the loop from because RHF is tuning for this so that you can automate everything. >> Yes. >> Did I miss anything else in the in the thesis of like what the northstar is >> there is that is that is right. I am overly nuanced in my communication. The one nuance is that we need to be practical. We need to be aware of what language models can do really well, you know, like what AI can do, right? Like there could be programmatic types that are like sick AF, but if you if the technology is not ready for it, it tots arrogant.

30:17 >> No, no, no. I strongly believe you. >> Cool. It sounds arrogant, but like I felt this way since long before I even had a company. >> I can I can vouch that >> I've been talking about for like 3 years. >> Yeah, I've been talking about this for so long and and I've been saying it because I thought it would have been easier. they say they do not do things because they they're easy. They because they thought it was easy. So something like that. I I thought it this whole project would take a week. and I was unbelievably wrong. So I am so sorry to everyone at OpenAI that I thought I was like man I'm solving this right now. but like I think that the tragic thing is when well I think overpromise underd deliver is tragic too and like AI is super extreme on that axis and I think RLVR is like the main well both RLVR and RHF are extreme perpetrators of this.

31:09 but like it to me it's like there's just so much potential there. Like AI is clearly so smart. I smart. I love this in my talks, you know, when I ask people like how can AI be so unbelievably smart? How can we like solve millennium prize problems in math but still not automate even the most basics of works? Like like really basic wrote stuff that like you know it doesn't take like extremely smart people to do this. It's not a satisfying job. Like there's other things these people could be doing, but yet we need them to do like this ba like super basic non unsatisfying stuff because you know like we can't automate it yet. But we have this like supercharged engine of automation that just does not have like the right plugs and stuff to plug into all of this economically valuable work. And you know like if the whole company of type safe disappears like maybe it'll take like a year or two for people to like truly catch up. I actually don't know how long it'll take. If if model quality matters, then we are going to be in a very good position for a long time, but it it like it it's done, right? Like there like this has changed the path of like technological history. Yeah. And like we will be exploring that space as a field.

32:22 >> Yeah, I I think I I definitely agree with that. you've created possibilities. So I think you know if I can paraphrase so that people can also understand you should not take the success of type safe and Jeff as just like well you know that is a new new model type now we're done we go back to business like no like actually there there are like five other model types that you should be exploring and like let a thousand flowers bloom >> absolutely like and some of some of which you will probably also early internet energy I think it's back to tech utopia you know it's no longer like oh man like sometimes my coding agents work, but that all of the best ones are hoarded internally, right? It's like creation is back on the menu. You know, though, it's going to be a wild ass world. And, you know, buckle up. I'm so so jazzed about that.

33:11 >> I mean, and now you have the the funding and the the momentum to do whatever you you you envision there. which I which I think is like very gratifying to see you have after, you know, so long of of saying these things, but like actually show the world. >> I know. It was such a such an interesting thing to be a tease the whole time. Like my talk like felt like it was a cliffhanger cuz I didn't say how the automation would occur. Yeah.

33:35 >> Sean reviewed our manifesto and he's like it's a little bit vague in these parts and you know like what's step one, what you know what is the intelligence. >> I asked you for model and you were like yeah model coming and like well I I just I mainly objected to the word composable but Bill Pron got is fantastic. Thank you. I I we we've really rallied around that. I'd like to think we're not entirely a cult like some companies are, but like we are like jazzed about what Butter doing and like we are like my brand is being practical and like we are all like so super duper practical. It's really great.

34:12 >> Yeah. So here and by the way here here's the the step the the secret master plan, right? The shape of machine native composable AI. It >> was your idea to make a secret master plan. >> It's a it's a Elon thing. when he started Tesla, he was like, "Here's what we'll do." >> Yeah. I I'm giving official credit to you. >> Thank you. Thank you. Thank you. but you know, you should have told me you're you're also going to do this model launch cuz you like you told you told me half of the story and then the other half you didn't have the Doom demo at the time. You didn't have any numbers to give me. I was like >> well the problem is I don't believe in benchmaxing. Exactly. Right. So like it is a thing that you need to feel and like I think that this is the way to build long-term trust even though it like hurt it hurt us a lot you know like like last year when we did fund raise no one believed us you know like like and they wanted just benchmarks and stuff and we're like we're not going to do that we are principled we're going to stand by our guns that rewards bad actors I don't give a you know like what do you want like this is who we are and we are standing by that so sorry Yeah. Well, in some ways I think like choosing the hard path, but you end up making the company that you want to work in. Y >> right. Otherwise, if you sell out, then you're just working in like Open AI but with my people, right? Which is >> Yeah. Yeah. I mean like I'm I I don't have too many regrets on that obviously like it worked out so unbelievably well and you know like I u I was emotional last night when I was talking about like the reasons I left OpenAI. and because like it actually had to change my wording after the launch. my my my phrasing was if an AI winter did happen and I did not do every possible thing I could to like avert that, I would see myself as personally responsible both for, you know, the the RHF direction, which I think really widened overpromise versus underdeliver and also not going all in on this because I think this is this is where value is going to just be like printed. So, and it was really cool because I feel like the the AI winter I'm worrying about is averted. You know, like AI will be useful. It'll be used for automation.

36:22 It's been less than a week and like the numbers are already undeniable that it's like being used for real work and like there's it's it's the wild west. Yeah. >> Yeah. Can you sh just just if you have top of your head, what numbers are you seeing like what's what's like signups like what whatever you can share? I'm actually not super on top of everything. Like the team is the ones who are telling me all of these things.

36:45 >> I'm sure it's like changing every day, right? >> It's it's kind of nuts. >> If if there's a milestone that you're like, well, yep, that's one thing we're hoping for. We reached it. >> I will say a milestone that we've passed is tokens per day. >> and this is not like fleeting tokens per day. This is like even at night like it's constantly churning. So, you know, machines are calling it and not just people trying things out. So that is that is so cool. A trillion tokens a day is a lot. so surpassing that is awesome. Signups to me don't really matter. And actually this was like a bit of a mistake we made if I'm like totally honest. People on Twitter were calling us like marketing geniuses and all of that. And that was just us. We don't have a marketer also hiring. and we were just being our genuine goofy like irreverent selves. And we were we were just like offboarding people off the wait list so hard. our platform team is so unbelievably cracked. I think we have more up nines of uptime than anthropic while having the most unprecedented launch ever. Like that is kind of nuts. So like props to them.

37:53 and the thing we didn't realize so number one weight lists weightless signups don't matter for like a developer platform in my opinion. you know, I would guess that a large number of them are not even developers. So they go in, they try some queries and a lot of people don't get it because they are not programming, right? Like they're just like, what? This is not a chatbot where where's my chat GPT2, right? But if like I I haven't exactly calculated this. My sense is that if every single human being in the world like just wrote a couple of queries, that would be a rounding error compared to like one power users for loop that is just like creating value. And the thing we didn't realize with a weight list is like we just off offboard anyone off the weight list. It doesn't matter. The scary part is rate limits. And then once people start getting value from that, then they just want tons and tons of rate limits because this is what software is, right? Like you spend effort up front to specify your roach task and then this wrote task creates more value than it takes to put in and then now that you have exactly yeah you run it in the background you make it a dependency to like other things you can make like higher level stuff and like you just create so much value in the world you know early internet people probably did not imagine like the wonder of early 2000s internet which is still not early internet but like it's it's through no offense composability that all of the crazy stuff happens.

39:19 I just really wanted to emphasize that in our manifesto. We are going for emergence. We are going for like being the catalyst. We're wanting to empower people and we are going to do whatever we can for that. Be it like Discords in our town hall with me wearing a garbage bag or not. >> and and podcasts and and you know getting getting like cuz I want the long form, right? It is like yes, we'll get past some of the superficial things and then we'll go deep and people will really trust and understand your mission and like you know the the people that will resonate that will end up joining you or or you know buying you no sorry as as a as a customer as a customer.

39:57 >> Yeah. Yeah, that was funny. I'm sorry. >> Sorry. I didn't I didn't mean to say that. U but no any one one version one very flattering version of this like 36 million views of your launch video. >> Cool. Up to 38 now. yeah, I'm rounding error. you know, Navia Stokes got 74, Fable 5 got 57. So like as far as and I I didn't I didn't do the stats for like original ChachiBT like which there was no video. Yep.

40:19 >> So like up there, right? Like as as far as as far as like if you were to launch a Neolab in 2026, I think you're like number one right now, which is like pretty crazy. >> Yeah. Well, I actually would rather I do have the shirt like your favorite Neol favorite Neolab's favorite Neolab. I don't give a about being a Neolab. I think being a Neolab actually we have a lot of like swag that's being a parody of a Neolab one of them one of them I have is like Neolab with product which actually is not a neolab like I don't care about that really what I care about is being a reliable dev platform so appreciate the comparison but like hopefully we transcend past them and we go back into like a thing you know like you know a revolutionary moment for developers and like this stable thing that people can rely on and trust.

41:04 >> Yes. I mean to that end I mean I I think that's one thing that really impressed me about you guys is is that yes you you you do talk about reliability. I thought it was mostly about calibration which like we you know we talk about RLCD but actually it's also about just like uptime and and scalability and all those things right they all sort of >> nines it's like like >> which time is in my >> but that's part of it but like there's reliability in like how intelligent the thing is like how consistently does it do the thing that you want and I think that like the the big reasoning models are very smart in my opinion they still lack reliability I think there's many use cases where you they look like they should be smart enough to automate their work. There is economic incentive to automate that work yet still they're not reliable enough as at an intern because they're optimized for different things. And so like I think that there's the reliability of being able to like trust the outputs and also we are like like there are dimensions of reliability that we are not yet at that I'm like so excited by you know like I want to automate the easy work before the hard work you know like I think that that's just a common sense thing to do. but to me we will be sufficient I don't know if there's a such thing as sufficiently reliable but I want to get so good that people don't even need to try the model to know that it'll work. It's like that's like what flow state is in programming, right? Like I'm just like writing queries because I need intelligence in here and you know like when for non-trivial branching I can just write it in in like a like like a type safe system one query and then get the results out and it just branches accurately like that would be so so good. Like that's that is the dream and that is like going to be like a long long slog.

42:45 >> Yeah. We're going to go into your API design in a little bit just just to give people examples and like maybe path not taken that kind of stuff. one thing up the front that I do wonder about in terms of reliability is I noticed that there's no seed there's no and and so basically same input do I always get the same output? >> If not why not? >> Oh great question. So this is actually like a common question we have between so reliability is actually a catch-all like whenever AI can't automate something it's due to some form of reliability could be like type safety it could be determinism it just could be like it's it's jagged right so reliability is a catch all I just think that it's also a catch all for like what the north star is determinism is like same inputs same outputs I do believe that this is slightly interesting for unit tests, but I believe that to be the wrong north star. I I believe robustness is what people I don't want to tell people what they really want cuz that would be a little arrogant of me. I believe that that is like the more important property.

43:53 you want given similar inputs get similar outputs. And it's kind of wild how unreliable LM are. Like a way that we test this is you put like uyu ids in you like little I think they're called nonses in in the prompt and what you want is similar outputs from all of those cuz it's truly semantically the same question and that is the part where you really want like like like that robustness is where like people get like burnt with AI making decisions. So I think that is the a super duper important property. We could also have determinism that is that is a thing that can be available as far as I can like mentally model for programmers like it it could be valuable for some use cases.

44:36 So like please educate me in comments or you but my in in general it's easy the determinism is something you can like trade off for better cost you know like we are we are constantly wanting to be on the intelligence per dollar frontier. We are doing like absolutely disgusting things to be there, you know, like this is I I shouldn't say this, but no one's here to stop me. >> you know, if you you sign off on your own PR, >> that is not how it works at this company. I believe for this week, my chief of staff, Kay, is the most powerful person in tech.

45:18 >> And shout out to K for organizing this. >> Holy holy She is so competent and powerful. She's incredible. I mean she sucks. Don't poach her. but so I tried to be a bit more filtered, but like people are telling me don't call it a Frankenstein's monster of models, but because that has like negative implications. I think Frankenstein's monster was like the good guy in this whole I mean it was innocent, right? I didn't read it. Okay, I I'll confess. Okay, that that what one facial expression I don't >> decent Jacob Vorti movie if you want to see >> I >> the adaptation anyway >> you have no idea how little time I have right now my priorities are sleep you know >> developers developers developers >> developers yes developers developers developers but yes we we do like absolutely disgusting things to be on the parto curve of intelligence per dollar and we are going to keep doing that we're gonna be doing crazy ass stuff. And I think people really need to think outside of the box like like part of the reason why surprising is like people are thought inside the box and we continue to do that. as of right now we are obviously the best at this and we want to continue being the best at that whole thing.

46:34 >> Yeah. >> So wait where did where did we tangent from? >> So so I asked you about will you have seeds in determinism and then you basically define reliability and like how you see it. Yes. But like determine like >> I have an robustness example that's that's real quick. I can show you. >> I love that. I'll just say one thing. We can make a deterministic model like like we're h if people can convince us that that is a valuable thing to do and we don't have a gigantic GPU shortage.

46:59 we can happily make all of these models we live to please. and re and and revolt revolute >> >> you'll throw over everything except you do it in a nice way and >> so so like determinism could be on the card. just gets you less intelligence per dollar. >> Yeah. well, just having seen the trajectory of OpenAI and topic you will just trust me now that you will be peer pressured into doing it. so like just people will want it. Even if they if you tell them they don't need it, they'll still want it. So like yeah, that that's the TLDDR of >> Okay. Okay. I will love to maybe one day we will see how that happens. I've been told I'm >> they they they say that part of our brand is being unshakable and they say that that's just the nice way of saying stubborn. Stubborn. Yeah. Exactly. And I'm a very stubborn person. I don't think we could have done that.

47:48 >> Yeah. No, but so like I Okay, but I tr I like have argued with you before. >> Yeah. >> And you've been right about developers every time. Okay. I give up. You win. You win. I I'm sold that I've argued with you before. >> No, no. I'm just saying like I I think that you can hold your ground while also like if I give you the right evidence, you can not you can sort like throw away your priors and be like, "Yep, like that actually makes sense to me." And so like, you know, just just trust your own gut on this.

48:13 >> I'll bring I suspect though that we will be GPU constrained for a very very long time. And anything that has less intelligence per dollar means it consumes more GPUs for the same intelligence which is you know like our goal is not to onboard companies like it's valuable but like our goal is to have people like experiment and do weird and we need like we need to like get it to as many hands as possible and like starting like the California gold rush for that.

48:43 >> I think there is right now. Yeah. just a word of caution. I mean I I'll just say it because somebody is thinking about it right now which is when you say things like we will not commit to deterministic models. We we will we'll do whatever it takes for intelligence per dollar and we are we are facing GPU constraint. people are thinking you may quantize your models, right? Like like whatever you had at launch, you may quantize down in to reduce the quality in order to to free up memory or bandwidth or whatever, right? And so you should probably have some kind of promise which you don't have to make now about like we will uphold model quality at launch people like when people when when we I mean you were at OpenAI when you launched all these all these APIs and even claude as well like when they first launched the models the model strings did not stay the same model at all times right you have versioning in your models that's great but like you should you should publicly commit to some kind of like once it thing is launched we don't change it >> we will not change our models when we deploy them. That is insane. We care about developers like like it makes sense if you're so doing something like that. Again, this is the problem with a for first party product and an API. It makes you can do whatever you want in a first party product, right? Like like more power to them. Whatever gets that experience, that is fine. With an API, you obviously can't do that. But I will say that we plan to move a lot faster than many people are used to model providers doing things. So we will be launching new models a lot faster than people think and we are not promising long-term support for the models because we think that there's lots of improvements to have. So there is a world that we might temporarily LTS what is right now Jev 1.13.0.

50:30 We might do that because so many people are using it and I know developers hate breaking dependencies. The alternative is fracturing our fleet and that is a very bad vibe for everyone. >> Have like 100 different versions of the model. >> Exactly. And if we're iterating very fast, there would be a lot of those versions as well. So we we do want to have not just a LTS supported thing eventually long-term support. we want a really sick way of doing that. we have like research stuff cooking in that direction and I think it's going to be the most prodeveloper thing ever.

51:04 but it is not yet our current models and I'm not promising that we will be able to keep the exact same models. They will get smarter every time for sure. And my sense is that even our model iterations where it already is smart it between model versions the the changes tend to be even smaller than the string models calling them twice. But when when it we go from like you know jagged to like wow that is where the big deltas are.

51:30 >> Yeah. one one thing one thing that's beautiful about LTS models is that actually you can also port them to other silicon. I don't know if you've thought about this. >> No comment. >> Okay. >> So I care about intelligence per dollar. >> Yes. >> Right. >> but speed. >> What >> speed as well? >> We'll see. >> Yeah. >> We'll see. >> I mean this is a whole part of the inference tech tree that is like I mean exploding in the past year, right? that you can move to a cerebrum and etched or whatever and get 100,000 time speed up.

52:01 >> Yeah. Like I I think that intelligence per second is like a different metric and we've even talked about like things like intelligence per dollar time second and like metrics like this. My guess on like Jeban's paradox occurring or at least the Jeb series of models. And the thing I like hunt people down about internally is like I don't care how much smarter it is. It needs to be in the pre-frontier. So like that is what the brand of Jev is. It is the best thing at intelligence per dollar for intelligence per second. We'll see. I I think that it's an intriguing thing. I know that there's many industries that are like extremely dependent on real-time stuff and they will like like intelligence per second means tons of dollars for them, but we'll see. I I I I would love to like do both and like have the market correct me either which way, you know, like like I I I I would love to be informed by people.

52:59 >> Yeah, totally. And it's not it's not just real about real time, right? It's also about scale because at at scale every microcond is just multiplied by billions and trillions of >> it depends on how background it's running right like if it's like a big background like database map produce query the latency might not matter so much as like the cost to get intelligence from it but like if it actually is >> something more realtime like userf facing you have budgets like between 100 milliseconds and 1 millond that are like totally magical and actually even if you were below 100 milliseconds if you half that time that means you can get double the intelligence or sequential intelligence calls to have like a like a phenomenal experience. So right that is definitely happening right now. It is super duper cool. I love the intelligence per second use cases but I don't think that that will be Jeb's niche.

53:50 >> Okay. Yeah. Fair enough. Okay. When think about promise of faster and cheaper typically the other the the trade-offs that other models are offering is faster but more expensive. Yep. >> Right. And so you're like one of the reasons I was thinking about why is Jeff resonating so much is that you've done the faster but cheaper side of the quadrant. Yeah. >> Which is very very unoccupied while holding intelligence like somewhat constant. >> yes. Yeah. Well I that that's a very loadbearing statement while holding intelligence constant. That's the hard part, right? Like >> which unfortunately like so basically you you refuse to have to like like do any public benchmarks or you don't like any public benchmarks about it >> and I will need some internal sense.

54:30 >> Say it again. >> You need some internal sense of this. >> Oh, of course. Of course. We have we have our own internal evals for sure. But it takes a lot of discipline not to game those and it needs to be like a top level priority to not game them. Of course we do that, right? Like how else can we make the guarantee that our models are in the prede intelligence per dollar, right? right? Like that we're not flying blind in there, right? If we're doing like completely weird things with different costs or whatever else, you know, like how do we compare them?

54:56 We plot them and get we try to figure out like what is the best for the users. So, we for sure measure them. I'm not anti-measuring. but it it it's extremely dangerous when you have like any alternative incentive. And this is the one thing that like I kind of rule with an iron well maybe my co-workers might think I rule many things with an iron fist but to me like not ourselves about how smart our model is is one of the most important things there like we need to be truth seeeking.

55:27 >> Yeah. Yeah. Agreed. Agreed. okay I wanted to go over some details on the API choices mostly because this is the only podcast that will ask you these these kind of questions. so you have three primitives. choice score no. First of all no where is that from >> is this like a term in the in the literature or what? >> now now it is we debated this a lot. We debated this a lot. it is you know it is boolish right like true false. it is >> but it's continuous.

56:02 >> Yes. Exactly. Oh so so first the origin of the name is Bernoli. >> >> yes so that's why it's even spelled that weird way that is like a subset of the name Bernoli from like a Bernoli probability right which is actually what that is. so that is the origin of it. We were debating this a lot. We liked peool we liked pool. We were we were wanting to call it like a pool party but then no one let me. you know, we had like a bunch of like other arguments about that and new we figured was like the best thing.

56:37 Our our rational and like this is actually the same thing with Jev too. is that we think that we are like an irreverent insane bunch and programmers don't care, you know, like if Jev is just going to be a string, we didn't expect it to catch on or even have puns or anything like that, right? Actually, there was a lot of hate on the name internally. They've all apologized except for one person. >> still holding strong.

57:03 >> Yes. our mutual friend. yes. Yes. Yes. Yes. >> I respect her for that. >> Yeah. Yeah. She wanted Jeb to be called Meow. >> She would, of course. >> Yes. Of course. >> Okay. You You win there. You win there. >> But but like Yeah. Reneual is we we had to make a new concept for this thing cuz if it was a bool it would be confusing to people. So actually all three of these are actually new concepts. These are not types that exist in programming and that was intentional because they map very closely to types but they're not quite that. A score is not an int. So if you had like instructor or paidantic or whatever map ints or floats into scores, you'd get a little bit cooked, you know, and and like we we were really airing on the side of clarity over the side of like making people like easily understand what's going on.

57:53 >> I mean, don't you worry about that. Don't you want things to integrate directly into things that people are already using? >> Yes. Yes, we do. And actually I think that you know >> you have integrations with like other SDKs and stuff but you have sorry you have your own SDKs >> but typically for example as a developer relations person I would be very obsessed with like yes here is how you use you know Jev with instructor here's how you know that kind of stuff >> we might have that somewhere I'm so behind on everything.

58:22 >> Someone will do it for you in the community. Not that not that you're successful people will be like oh that's cool >> but like you know >> I don't see that as binary either. I actually see success as a score and there's always more to climb in like how much we can like be there for our community just to be clear and I'm this this section is stressful cuz I didn't review the docs and they're constantly changing but to me scores do exist. So so so scores are similar to like LM judging, right? So like if you want to call it like a judgment I guess you could but like that that is like the the the way people already use this type of thing right like maybe a new could be like a probability but everything for us is a probability and a choice is actually closest to a function call but a function call is like an extremely disgusting thing that if you want open AI juice sauce tea that we should go back into that later like Like a choice is just like the right way of explo of exposing like a switch match statement. Yeah. Within.

59:26 >> So so it like maps cleanly to an enum. Yep. >> And you can choose to hydrate it into a function if you want. >> Yes. And like the enum choice is the important part of that. And like actually I think these map all to like programming primitives where like choice maps into like a like a switch statement on an enum. nules map to if statements and scores map to sorting or thresholding at a greater than or less than. Okay. And this has been always what the vision is like there will be more types and they will map into programming primitives.

59:57 >> Yeah. any other so any nuance you want to go through for literally this is for the Jev people who are like deciding to really invest in Jev you are the expert right? I'm just like wanting to provide more background for them on API choices. you know how they should use some of these these things like legends confidence how critical in your testing you know like how how like just any sort of like pro tips that you want to offer people when you're down at this level.

60:27 >> Thank you. I love this. no no this is why we're here. >> Hell yeah. I didn't expect this and actually I no one has asked me this in probably like months when I was like onboarding like our devril. Okay. >> so sick. so our model is designed for being like deep in the insides of computer programs in the future. We like unironically believe that this will be much more massive than anything people are even considering today. And our model might not be ready for that but we are like continuously working for that future. it will never be good enough at these shallow tasks. sorry, it'll never be like we're not just going to keep on climbing the shallow tasks. We want to be deep in the guts of programs because that's how you make software powerful. All the all the types inside of our this is actually an output.

61:16 but all the all the parts of like the the input like the state, the instructions, the criteria, all of them can be structured JSON objects. That way like programs can like insert them in the right spot and you don't need to like put things into templates. Exactly. So if ever I I think people don't read into this part enough and they think it's all strings and that's that that's fine. but these are all meant like like I would say that if you're using like a template like turning it into like a system message or something, you are thinking in like the old way, you know, we should be making things as easy for computers to understand because that structure is truly there, right? Like it would be weird in like like a programming language to have like all of your numbers in and then you pass it into like you turn it into a string.

62:06 Normally you do that for printing when when you have a human in the loop, right? But for like within the computer, you want to be passing like nested structure that is semantic all around. And we are really going to be optimizing our model. The model's pretty optimized for this. But the thing is every different nested level of structure is harder to reason about. And we want we are really cooking hard in that direction. I think people should keep cooking that direction because it makes the code like so much more legible and beautiful and like agnostic to like the the implementation details. It's like here is my state, you know, like here's my function state. Like think of think of it as like an AI function. Which subsets of my state, which is like all the variables you have available, should I pass in here? System messages are like disgusting global variables where you just put everything in there and you put all the instructions at once. And then you know like you you hope that every single instruction gets nailed instead of asking the questions in parallel.

62:58 >> Okay. >> And also I would recommend I and I truly say this not from like a like it makes me money perspective. I truly recommend asking lots and lots of questions, break them down, make them smaller and like really decompose like no matter if the models can do it today or not. I believe that the biggest like saving grace of like what's happening this week will be people's code bases, AI code bases are going to be so much better. You know like if you decompose problems into simple decisions every single one of these things is extremely evaluable like like a AI beforehand is big system message and then maybe you have like another big AI to see like if it actually does this. That's nuts, you know, it's kind of crazy. Like it's that was like Stockholm syndrome, right? But like that's kind of crazy. Like if you want to say like, hey, don't read this subdirectory or don't pass any API keys to deepsek or whatever else like that should be programmatically basically guaranteed. And you know, you'll never have guarantees of any machine learning model, but like by breaking it down, you can actually you can actually measure it, right? you can verify that it was actually called >> and like our model our model like the the interface itself is so verifiable this should be like a sigh of relief you know like it's it's a it's just going to lead to way better engineering >> yep I think I get that and and so you know one of the reasons people didn't used to do this in the past was because they would just call a a small LLM right and it's still too slow it's still too expensive versus chunking everything there I've done exactly this myself right like I I benchmark here's a pipeline that fills everything in system prompt and it just gets one big output versus break it down into 100 different things. It was slower, more expensive, not as good.

64:42 >> Yep. Yep. Yep. And that happens. Yeah. It's and it's like super inconvenient. It's unwieldy. Why not just put it all together? You kind of end up repeating some stuff between questions. So, it's like maybe like, you know, inefficient or something like that. But then it results in something that is very hard to rely on. And software >> doesn't need to run in the background. It would break my heart if our stuff couldn't run in the background.

65:06 >> Is there a way to break things down that you guys have found that works versus what you thought worked and doesn't work. >> Interesting. >> Because like people are just going to be exploring this, you know, now that you've said it. Like they would use this as a reference and be like, "Okay, like that's how I'm supposed to use Jeff." >> Yep. >> then the question is how do you break things down? >> Interesting. I I like to break things down into its like its smallest semantic unit. Like what is the lowest level thing? I try to never have I've probably queried the the model the most among anyone and like I try to number one in my in my queries this is this is a lot more like the way I prompt things like I make it really really structured and explicit and in the questions I always I like the back ticks but like it works for all of them you know like be really clear what I'm referring to because we want the model to be really literal because when you program you want things that instruction follow really really well. That is what the art of programming is and what AI does is expanding the things the kinds of instructions that can be followed. So I'm a fan of doing that. I I sometimes I'm a little lazy and I like I have like more like hybrid things, but like I think that for like really big production things, you just want to like keep on adding more questions. You want to make it really easy to add more questions. you know, be really really precise about all of that breakdown and then have the code to have the exact behavior you want. If I could give like a tiny little example of this is like refusals, right? I'm not going to talk about why we don't refuse.

66:43 I might have done that already. It's like all blur. but like for refusals, I don't think you should ask should I refuse here, you know? that's a really I think the answer will be pretty good because like that's a system one compatible task. But I think you're way better off like asking many different independent questions about like the different situations you can refuse about because instead of having to like just guess based on you, you know, you can actually specify what you want and beautifully, you know, and I think this is like truly really beautiful. If you find a situation where it's like, oh, it didn't refuse because of this reason. I didn't specify this part of the task. That is awesome.

67:20 That's what software engineering is about. Like you fix the bug by adding that question in, adding the threshold, maybe remembering that as a test case, and now it is just solved forever. Like your software can't forget about that like in the prompt because of context rot. It is just there and you can like just keep measuring that forever, you know. And if the models are not perfect at some of these things, you can choose what threshold you want for all of these factors based on real examples. It's like it's like ML without the ML. And you can just do it for for anything. And like there might be some things the model's not good enough yet, right? Like I would I'm a little bit afraid when I see people doing trading with the models like automated trading.

67:59 >> it it looks cool. I I I just think that people should leave it to the professionals. and like that's just a very hard highlevel task that maybe the models aren't good enough yet to figure out. like even if they were then they would it suddenly wouldn't be because of efficient market. But like that's one of those things where you can like break it down into things and just evaluate them and you might be like it's not smart enough at this. Maybe we don't deploy it yet for this version or we make a trade-off or we air on the side of safety or like hey the models are not good enough at you know like detecting like this weird combination of like sarcasm with a VIP customer that this is when we escalate to a human and that's what confidence estimates are about too.

68:40 >> Okay. great great answer. I think one thing I I'll mention very quickly which I don't expect that you have a too long of an answer for is well no no no it's just it's just specifically like you are still relying on thresholding as like the the lever that the user can pull but what if just the the the calibration is wrong right like you're saying your your calibration is perfect but >> I didn't say that I didn't say that >> so perfect calibration means like lower value is lower like so probability lower higher value is probability higher but it could be wrong. It could be locally misaligned and so then I would want to fine-tune it or something right which you don't offer but you could again see this is a short answer which is you don't have it right now. Oh, do we want to offer fine-tuning is it the question?

69:25 >> That could be that could be one version of it or you could have a different knob, right? where like because like right now you're all you're saying is like if something's wrong skill issue, you should you should just change the prompt again or break it down even further or you change the confidence. Those are my two options, right? And that doesn't feel super satisfying if your model is just getting it wrong. >> Yep. And it will it will get many things wrong to be clear. We have like a report issues button. Complain to us in Discord. We want to make it a lot better. every single model version will be like noticeably better. We will stop shipping them quickly if they weren't getting big improvements. So number one, that is like totally reasonable. I think that that's simply pragmatic to admit that AI is imperfect at some stuff, right? I do think we'll find use cases that they are like good enough at and good enough kind of depends on the use case, right? like human beings can do a lot of work despite being bad at that work because their EV is quite high and presumably with the right thresholding and everything there probably is like large amounts of work that could be done even if mistakes are being made. on the question of fine-tuning I could imagine I could imagine it in the cards.

70:31 I do have concerns because like in the what people need versus what people want category. like I think general models tend to be really like like again there's the geniqua of generality that making it good at like a million other tasks than this one narrow task might make it better at edge cases in that task which I'm I would be a little bit afraid of you know. Yeah. I I I could imagine it is is my answer. I'm endlessly practical on these things. I want everything like my vision of the world is I there's there's so much we want to be building but also like I would not want to ship something that is like a giant foot gun like some other AI companies would >> well know so both open and claw and and I think even Gemini have rolled out fine tuning and then took it back y which is an interesting observation that pretty much fine tuning is now in the domain of open source models Yes. Yes. I do know about that and like it was kind of crap. So like that's probably better that they took it.

71:39 >> It could just be a foot gun and and telling people that fine-tuning it is probably the wrong way to go is great. Another interesting answer could be that like well our model is so different like you know in the same way that quantization doesn't apply to us. Output tokens doesn't apply to us. Fine tuning also doesn't apply to us. >> Well actually I I'm I'm super open to that possibility. my this is not a promise. This is a desire just so to make it clear I like to be really honest. Like I think that as intelligence per dollar gets cheaper cheaper cheaper cheaper I think that we could get really like small approximate things that hopefully are proxies for intelligence. Like is there a world where people don't write reaxes anymore because like you know the intelligence per dollar that uses AI is cheaper than like the complexity of a reax you know that would be kind of sick. I would love that, you know, and it might require fine-tuning for some of those narrow use cases to really get past the threshold.

72:35 We will see. My hope is calibration gets that. Calibration plus a cascade of models, like if it's super constant, then maybe it's right. And and if it's in the middle, then you do the next bigger model and you chain off from there. I don't really know how that's going to go, but yeah, I I I could imagine it. And something that I could imagine too is like imagine you have like a a series of like we own the entire per frontier something that a business might want to do or I I think a hacker would be okay with dealing with a parto frontier of models. Maybe a business wants something more dynamic.

73:07 You could imagine like having like a different sizes of models and to dynamically pick which model based on how smart it is on different parts of your stack. And you could even imagine because of how simple our thing is, you could imagine like some automatic fine-tuning on that. >> Yeah. Not a promise in the slightest. I'm just like cooking on sci-fi. >> But you would consider different sizes of Jeff models so to offer that. >> Absolutely. Yeah. Yeah. Yeah. Like we we like how would I know how much intelligence people need, >> right? Yeah. I don't know either.

73:37 >> Demand is demand is unlimited. >> Well, yeah. People are telling us not to ship things right now because we don't need to ship things because again >> it's good enough. Yeah. >> Yeah. But that's kind of lame. And I really like the saying this is something that I hope people hold me to because it'll be hard to to to walk back from. Yeah. Like that I I don't know if exactly the thing that culture is what you do and the market doesn't reward it.

74:03 And I really like that because I think that we are standing for something. Maybe in the future what we're standing for is like so obvious that we're the equivalent of like boring like Visa or something like that and like we're just like a utility that no one really thinks about and I'll be wearing non-pink suits or whatever else. But I really want to be like rallying the world to this, you know, like I want to keep doing cool stuff, not because we need to, but because I want like people to realize that this is just the beginning, you know, like that wasn't even meant to be the opening salvo. That was like kind of like a, you know, low-key research preview or whatever you want to call it.

74:44 And there there's a lot more we can do with with like machine native intelligence is going to go wild. So not the only potentially not the only size potentially not the only model that you guys launch you know that you >> absolutely not for any of those I want >> I want to like meet whatever needs we can right like at but with like a giant caveat I don't want to be like open product teams that like like throw stuff at the walls like I want it to be like in under a unified vision like if you go back to the manifesto like everything needs to be under one of these three three things in my opinion >> >> I'm Yeah, >> I'm not prepared to do this.

75:21 >> Oh, I'm sorry. I'm sorry for asking. I can just just talk about it. Like we have like three steps in our stuff. It sounds like a tease. I want everything to go under one of these three things to keep pushing the boundaries and everything like like this is not these are not like checklists. These are like axes that we think build like the foundation of, you know, of like a new technological revolution. And I want all of the all the bets we make to be somewhere in there. and we will be doing some weird weird stuff modelwise. So, because machine native, right? Like viewans don't need to totally get it. It needs to just be valuable.

75:57 >> you know, with just give people a tease or a hint like what what does weird look like? What is weird? >> I I'll give people a hint. Yeah. >> some people are trying to call them decision models. >> Okay. >> the our primitives are decisions. I I wouldn't do that. because I think there's other types that are machine native that are not decisions. >> Okay, we'll leave it at that. I think it's I think it's a pretty fun hint.

76:27 >> Yeah. Yeah. There's people, look, let's there's people saying like, I've done this before. I made a decision model a year ago. Like, Jeff is not new. Jeff is not cool. But like I I think there's the categorical like here's what you're establishing is possible. There's the performance of like well actually the for the benchmarks and the numbers that you're getting you are still beating as far as you can tell you're still bidding every single clone of you out there.

76:48 >> I I don't care about the benchmarks just to be clear. So like like even if we were winning or losing I want to deny >> you establish the category right >> but but also I think this this nuance between decision models and system one I think is actually the thing that you're trying to >> Yes. And I just want to make software engineers superpowered, right? like with AI like and or like the tragic thing to me is, you know, in that AI winter direction. I think like it's it's it is so sad that AI was so powerful yet so underutilized.

77:26 it's the thing that gets me emotional. But man, like I think that that is I don't want to like just be like pure techno optimist like all technology is good. I think what was happening now was like a travesty. Like it's and like there's you know I just want to like open up those possibilities for people. Yeah, I'll just end it there. I I I I I've I've cried too much these last few days to to to want to do it on the record.

77:55 >> Yeah. Yeah. No, I appreciate you sharing a little bit of that and I think people can see that you're very authentic and passionate about this. you know you don't necessarily get that from the name like type safe AI but like I think once people immerse themselves enough in like here's the genuinely different direction you want the world to go and like actually you have done like the the hard part about going zero to one on the on the thing then like like now let's all go to go together in like the new direction right >> yeah yeah I I don't I am sure that I won't think maybe I will think that the hard part was done perhaps I think that there's going to be many more hard parts like if you know all sorts of stuff gets automated and we finally see GDP growth and like you know it's like you know a JF party every day then maybe the hard part is done but like I I don't think so and like I really really think that people focus too much on speed and cost and not enough on reliability like reliability is what makes it delightful like reliability is what like allows you to trust it >> you have this line >> TFP grove reading 3% % in 5 years. Hell >> yeah. I've never seen >> Hell yeah. Let's go.

79:04 >> I've never seen a lab care about TFP grow. >> But like that is what an economic revolution is, right? Like it's actually extremely consistent with what the OpenAI charter used to stand for. You know, it was talking about like I think the charter is the same, but they've kind of tried to move definitions around to like, you know, 100 billion in profit or something like that. Not that I hate an open. >> It wasn't like Yeah, it wasn't a well- definfined term what AGI is, right?

79:26 >> They tried to do it, right? like doing majority of the world's economically valuable work and they should have to answer the question how can it do millennium prize problems in math and zero of the world's economically valuable work rounding error you know like I think that all models are roughly tied right now at zero there's some chance that like we have started already but like I would guess that it's not yet 1%. and I think that that will show up in like when it does happen, it will show up in the economic statistics. It's going to be awesome. It will not cause mass unemployment, but it will cause like a whole bunch of awesome shifts and the world will be a lot better. And also like I'm really tired of AI always being the foreground character of things. Like I think that the world should just be more delightful and AI should just help with that, you know? And I >> just like disappear into the background.

80:19 >> Exactly. you know like the I I say this in my talks like how can it be that 2019 software like software SAS whatever super duper valuable right it's 2026 now how is the software basically exactly the same despite AI being so freaking awesome other than sometimes having a chat box on the side right that that like that kind of works but doesn't allow you to make decisions that the companies have stakes in because they can't be trusted to make decisions that to is nuts. You know, there's so much economic incentive for this and I I think it's going to be like like an inverse SAS apocalypse. I think SAS is going to be supercharged by this. They are the ones who are like most in the know of what things are valuable to automate and it's going to be like a crazy time.

81:06 >> Yeah, I I think I think so too. It's a it's a beautiful thing that you've unlocked, you know. Yeah. you mentioned one thing here which I don't know if it's like directly here which is what is a system one problem and what is not what is a system two problem like you know >> that's a hard one that's a hard one my friend >> cuz people now are just trying to jev everything right which like probably is going to fail right but like some things are going to be good >> je everything is a pretty funny it's a pretty funny way of doing it saying it the so I'll tell you the truth >> the truth is that this is an empirical problem just like scaling laws are an empirical thing you know like why doesn't like robotics really work right now despite all the money being spent on it I don't think it's about like spending more money necessarily the the empirical results just might not be there right so empirically I believe that these like pre-trained super condensations of intelligence are fundamentally system one thinkers I think that they truly like system one is the closest thing to describe what LLMs are strong at. RLVR has done incredible things for system 2 thinking. I am at awe. It is super freaking cool. Like I don't think that it's going to result in AI doom in the slightest. not not 0% of course cuz I think 0% is miscalibrated, but like it's it's really cool what they've done and they've really pushed it to the limits. Well, maybe they don't think so. not the limits limits but like it is it is a weird thing for models to do and they are very fragile at this like think about how people used to talk about AI back in the chat GPTJs like wow it's really general it can do a lot of general things and and and then but it's bad at math problems and like GSM AK grade school math and then now look at how people talk about RLVR it's so fragile it's so jagged you know like it can why can it do this like really weird thing and actually you know math is not just spiky it's fractal right and this is because RLVR is you know like if we talk about like what is the north star for each thing RHF is please humans right that is what the human feedback is RLVR is optimized benchmarks you know that everything that goes into the RLVR category literally is a benchmark by definition because a benchmark is programmatically verifiable simple outputs that can like do well and you know RLCD is make it reliable for you know programmatic use and yeah that I'll >> yeah that's maybe I I'll offer some thoughts and then you can sort of correct me if I'm wrong one for example one thing that I've been thinking about is also so I threw Jeff at a bunch of things when you gave me access on day one and multihop reasoning right like so single hop fantastic like state-ofthe-art you should never use anything other than Jeff for single hop Yep.

84:01 >> Multihop is going to start to falls down and it's like kind of monotically increasing as you increase the hops. >> Yep. Yep. Yep. >> Right. >> So, oh yes, back to that empirical question, it depends on what we can like pull out of the models. Right. So, we want everything like we want to unearth as much intelligence as possible period. the models like I see us as like unlocking and smoothing and sculpting the intelligence while like adding new capabilities and like you know like like filling in gaps in it. and we will be filling in like you know more and more and more and more of these gaps over time. But the reality is that we are in the business of unearthing properties.

84:40 Those properties are actually a function of what is available from like these like you know these condensed cores and like franken signing them all together to have all of the properties of everything you know. but the reality is we are in the business of unearthing as many capabilities as pos as possible and system one just happens to be the description of what works and everything that works in that paradigm will be system oneish you know like I'm like like there is a reason why we don't do what's called latent reasoning reasoning in strings I think the reasoning like what models do really well is reasoning within the models it's not totally complete it doesn't do great at all >> latent reasoning is reasoning in strings I thought l reasoning is reasoning in in inside the model weights >> that >> I think that people used to call that continuous reasoning. I'm not entirely sure. It was called latent reasoning because like it used to be that the reasoning traces were secret. So they're kind of like a latent variable for the answer.

85:36 >> Yeah. >> So what's secret has now shifted. >> Well, it's still secret for open anthropic, right? >> So no reasoning Jeb. >> Yes. >> As far as you will ever do it, right? Because that that like violates the whole promise of system one. I my promise is to do whatever necessary for machine native stuff. I I could imagine there are there are some forms of reasoning that are less slow, inefficient and fragile that I that are like totally on the cards just to be clear. So pragmatic person, I'm not making promises on it like you know like methods. I'm making promises on like the what my ROI northstar is and I'm going to fight for that like like you know like like this launch didn't happen and we are still like hungry for our place in the world.

86:22 >> That's great. Yeah. >> Yeah. >> I think the other thing that vision is another one that's like a big like you know cap capability that you don't have but maybe it doesn't ever belong in C system one. >> I I think I have a pretty good vision. >> what? Sorry. >> I think I have a good vision. >> No no no I'm kidding. I'm kidding. >> Oh my god. >> Yeah. Yeah. Yeah. >> because people obviously the first thing they want is vision because of the Doom demo but also just like everything you know other than text is vision.

86:48 >> Everything is in the cards in my mind like like and actually this is like a debate who we have this man your audience is probably like the great one to have in this debate. There's a question about like how much do we try to like give people what they think they want which is what we did in stealth for 2 years. we just knew that this is obviously going to be valuable versus give them what they say they want, right? And like there's a lot of dimensions of this, right? And you know like context length is an example of this, right? every single model including ours, I actually think as far as I can tell ours is like by far the best at not degrading in long context. but like the other providers are just like whatever people wanted, let's just give them the stupid thing.

87:34 And like we need to figure out a balance for this because you know like if you take the the the former side too far, give people what they want, you end up with like anthropic nanny state style thinking which is very like anti-developer while like the pro-developer route would be like give them what they want but developers are like we don't want to put the burden on them to figure out the genes of intelligence. So we are trying to like figure out this navigation of like how quickly to release things to still like have our like brand of trust and also like teach our treat our users like adults that can make informed decisions that don't need like nanny stating on top of this stuff.

88:15 >> Yeah, I think that's fair. >> And and we don't know the answer to be honest. we'll we'll have to figure it out. It's going to be that's probably going to be like one of my biggest debates over the next couple of days. Yeah. because like we have a lot of stuff. Again, we didn't expect it to pop off. So, we were like, we we'll need some follow-up launches. >> Yeah. I don't know. I don't know if you didn't expect it to pop off. Like I you I saw the work that you put in. Like I have never seen you lock in so hard as like the last two months basically, right? Like >> Well, that's also because my chief of staff made me lock in.

88:47 >> Yeah. Like it's like I have never >> I thought I worked hard before. >> Yeah. And >> no, but like you were showing up at our writing workshops and I was like, "What are you doing here?" And and like it was useful. It was great. >> You clearly like were very intentional about your launch and the work showed and like congrats. Like you know, I hope to >> keep locking in is my is my sense. I I want to >> like like I think that we've passed many great filters for the the the tech world that we're wanting, but like there's still going to be a bunch more. And like, holy smokes, am I excited to fight the good fight.

89:25 >> Yeah, it's exciting. before we broaden out to topics outside of Typesafe, I just wanted to offer any other things that you think like underrated or misunderstood about what you have launched. >> Underrated or misunderstood? >> Yeah, you have panouts, sorry, patterns here. maybe maybe want to go into that. model jaggedness, any anything, you know, >> give me one noodling of it. Oh man, I I would rant about all of these. I I really shouldn't. I really shouldn't.

89:59 >> and people can come go to your discord if they >> People put a lot of love into the cookbooks is what I will say. The cookbooks have like some fire stuff. We had considered putting a bunch of these things like in the main launch blog post, but it got kind of long and unwieldy and like very power usery. But like we really really I'll be frank like before the launch every like what we're saying sounds like sounds like this weird alien tool. Why would anyone need this? You know it was a very weird thing and we were very worried about teaching people about like this new frontier. It obviously succeeded but like we put a lot of work because we thought that education would be like a gigantic bottleneck for us. I it probably works and it's no probably no longer a problem because people are doing things like well beyond what we could ever show you how to use your model.

90:50 >> Exactly. But like they Yeah. And their use cases are like kind of cooler than ours. Like like there's a bunch of stuff where I'm like, man, if that was our demo, holy that was way cooler than what we were showing. like the computer use stuff. Holy smokes is it cool. but like like we put a lot of love into this. This is not like AI generated trash as far as I know. We put a lot like it's like a lot of love in here and like each of these are like like there's real alpha there like these are inspired by solving real customer problems that existed and we went through the work of like helping them do cool ass stuff.

91:27 >> Yeah. How much while you're talking about this right how much validation did you do before launch? Like what you know what was that process like? What was that process like? >> Like clearly you did some but obviously you're not getting in touch with as many people as you are today. >> Yes, of course. I actually think that the reception was pretty bad and like actually for the non-technical people in the team they were really worried >> you know like there was a lot of fear.

91:55 It's like no one really gets this and like you know they don't want it. We're like selling like a vitamin and not like a painkiller. like should we have FTEEs to like write the software around solving that problem? We had almost no revenue before launch. it was kind of like we like the technical people were like obviously true believers, right? Like we knew that this was sick. It's prop computational properties are like off the charts on like so many axis that we're like yeah obviously it's going to be huge. I was definitely super afraid, which is why I locked in super hard. But like the most common thing was like I would say like more than half the people we had play with it just did not get it. And like the people who did like were like man this is really cool but how do we get this through procurement and stuff like that you know just like it like quite a quite a battle and we just knew like okay our target market is going to be developers. people will find the use cases and that way everyone is gonna FOMO in and like I I don't want to rub in people like changing their minds with the facts changing. I do want to call into question like the concept of product market fit you know but like because like there was a product there was a market like we were like hey do you want to use this and people are like I don't know really know if it solves our problems it explodes and everyone's like we need as much rate limits as we can can we literally give you GPUs because we are constrained right now so of course you know like marketing is an element of it of course but I don't even think it's about marketing I think it's about like passion developers who've like you know our souls basically resonated at the same frequency and that frequency and that got got everyone else excited too and I'm hoping as well that like we as a company will be eternally eternally eternally grateful to those developers like not and not just like you know like the the companies that like are like start off with developers and like go to enterprises. Exactly. And like I'm like even thinking about like how can we launch things that are better for oh man I don't know if I should say this but I will >> better for developers than enterprises.

94:10 >> Exactly. How do we do that? Like how do we empower them? and I have cooks. I have cooks but it's a it's a very weird thing to do and like I don't know how else I can show my thanks and loyalty to that you know like and and that's why I did like the the dying my hair yesterday. It's like it's like I wanted to talk to them cuz it felt dirty to me during our company's like most important times not to keep talking to them.

94:36 >> Good. Well, I mean that's why one of the reason you're here >> hold me to that please. I I try to be principled. Quote me on this. Call me out. You have the pitchforks out if I change. >> I was just going to briefly show the computer use stuff. Is this is this >> I've never seen I haven't I've seen I saw like a airline browser used thing. And inside this new note, let's make the title say hello.

95:02 >> Wow. >> Great. Great. Okay. let's move on. And can you open up the Arc browser? And once you're there, can you Google search Norbert? >> now, can you open up X.com? >> Is this this kind of use case? >> Oh, the voice use cases. This is actually the first one I've seen. This is >> Open up the photo. >> Wow. >> Oh, wait, wait, wait, wait. Oh, can you can you go back a second? Can you go back a second?

95:27 >> rumors claim anthropic engineers worship Claude as God. Wow. >> Wow. Dang, that's pretty funny. >> and here you are building prod. Wow. This is sick. >> yeah. So, so clearly you can operate the whole computer with voice with Jev as a decision model. >> So, just like I'm anti-benchmaxing, I'm also anti-demos. I want to make sure that it works reliably. I love people are playing with it. This is super sick. Have no doubt. I want to see this. I want to see it be used. I want our team to play with it. I want to find the weaknesses and I want to solve that. And I would Man, that looked really cool. That looked really cool. I want that. I want that. Like when my when my wrists are sore, I just whisper flow everything. That would be sick.

96:13 >> Well, as you know, just just to round out the use cases side because I I do have to let you go. who's, who are the who are the bigger companies that have reached out and have surprised you with what they want to do. >> Just >> I'm so out of touch for that. People have shown me screenshots of companies and from what I've seen, it's all of them. >> Well, mostly like, you know, for for those people who work at larger companies and they're not doing this kind of work, I just want to give people examples of like, you should go look that up. Look that up, look that up.

96:42 >> Oh, so like I think demos are super duper sick. Obviously the coding agents are like gigantic use cases like they are like also super sick. >> Call is all about Jeff right now. >> Oh hell yeah. Oh can I can I give a little bit of a a tangent about coding agents if that's okay? >> Yes please. >> Oh let me give me a second. Give me a second. Okay actually I'll come back to coding agents. Let me describe like the big families of use cases. Like we've mapped this out from first principles like long before release. They are what we call dark data. Like people hoarded big data but they would not throw LMS at it because it just was too expensive. So large companies adore this. They have like piles of data that they wish they could analyze and this is like a data scientist's wet dream. So this is like this is a giant one. Like I think this plus coding agents are the big money makers because that's what the all where all the volume is, right? there's the real time stuff, you know, like people who need like intelligence in the loop.

97:40 they like I would guess that every CEO if not CTO at those companies knows how much better their product gets with every like 10 milliseconds shaved and like >> especially e-commerce. Yeah. >> Yeah. Oh, or like assistanty things, you know, there's many AI assistanty things and like as far as I can tell, they really love it. again, I'm not in the front lines of customers right now, so I just get know what my team tells me, but like this I'm so excited for this. I'm really excited for this for games. I really want to play like sick ass auto battlers where you're like commanding your team or like semi-auto battlers. I I think that'd be so cool but don't make it too good while I still have a job.

98:20 and the you know like there's the the the what we call like verify everything you know like verifying all LLM calls kind of like observability. I think actually on the note of docs what people should be doing is like the parallel questions are very cheap. So if you have like big states you want to ask many questions on >> right here. Yeah. put ids on every like message and then ask a question about each ID. So like when you have like a long state >> so that way you can like pay for that state once and ask lots and lots of questions about each message within it.

98:50 I think that is like a like a great way that like saves money and is >> by the way I always think like it's interesting framing system one and system two because it basically makes the case that you should always make one or 10 or 100 Jeff calls for every one reasoning call that you make. Well, maybe. Well, I I mean, I don't I would like people to spend less, you know. Maybe you do like, you know, one half the reasoning calls and like 10 Jev calls each or something like that or whatever solves the problem that like couldn't have existed otherwise.

99:23 wait. And number four use case was what I described as like smart software, like software that's intrinsically composable and like does like weird fun stuff that could never happen before. you know, like the programming language as Jev thing. I don't know if you've seen that. That is so cool, man. If we knew how to give out credits because we're really early in our infra days, I would want to give all these projects credits. >> and I think that those are like how we've mapped out like the main use cases. computer use has also come in kind of like the real time direction as well and like that's really really cool.

99:58 If it is reliable, I am super jazzed about that. I suspect we can make the model a lot better at these use cases because like that came out of left field a little bit. So that that's really cool. on the coding agent thing and this is like a really surprising thing that is happening right now. Okay. >> cloud code and codeex are I believe the the winner like the the number one and two. I'm not entirely sure. I don't follow closely but like it's roughly that. But they're built around a single model world, you know, like and and that makes a lot of sense for them, right?

100:33 Because like it has been a one model game where it's like kind of like the same model but different intelligence that you're shopping. But all the open coding agents are like jazzed right now because they're like getting their Jevon and like the thing is there's I'm sure they're trying a lot of weird stuff. >> Mhm. >> But all the coding agents are kind of roughly at like approximate par, right? Because like there's not so much you can do with a Y loop. But the moment one person finds one killer use case that you know you can only do with that coding agent everyone will flock to it because they have like a monopoly on that thing but all the open coding agents will be able to copy that >> right and but I don't know what the cloud codes and codeexes will do because they are built around that one model world and like I think that's going to be like a really interesting thing you know like I would love to be able to integrate with them personally like I want to integrate with everyone like I they might make competitors eventually. I don't know. But like it is not me my job as Sonfire Infrastructure to be opinionated on that, right? Like I want to just serve the world. but I I don't know if they would do that. And like I think it'll make the coding agent game super weird, you know? Like I'm so excited for that. And like I'm sure I'm getting my team to review right now an internal document I made on design patterns I suspect will be useful for coding agents. So hopefully I can share it like right after I walk home.

101:58 but like I think that there's just like such ripe area for exploration out in the world and like it's it's man if I did not have this I would love to experiment with coding agents right now. >> Yeah I mean and I'm sure the coding agent companies would love to work with you as well to to figure that out. yeah I I do think that there's still use cases for cloud coding code with you guys which it's it's easy to explore there. okay I mean you know we've you've you've been very obliging in the sort of indulging in all these all these things. I just want to take you out of types safe just generally about and you've you've made very clear your position on on the state of the eye. give you more room on the alignment safety side of things.

102:37 >> Oh did I not talk about safety alignment at all? I I think I didn't I think maybe I didn't. >> You did. I just like you know I I think that there's there's a lot of you have a lot of researcher discussions. We have this every every in Europe. What are people talking about? You know, like I so for example I recently was at one of these researcher gatherings and people are genuinely worried about the pacing, right? Like this this whole topic about like we should slow down because the public is like clearly not ready. and I I'm sure you have strong feelings.

103:16 I feel like this is the kind of thing that is a dangerous topic to talk about. I'm happy to talk about it. I I live for danger. >> Our company brand is chaos. It's not Jev. It is irreverence and chaos >> and you know. Yeah. And like you were at OpenAI during like the one of the very first like very visible incidents which is the blip, right? like which like and like the dominoes have gone down now to now every Frontier Lab has co-signed a document saying that they want to base >> interesting I so comp it's a very complicated nuance thing I actually do want to write a response to this more formally I do have like a little bit of a short version of my response yeah >> which is that as you RLVR more like RLVR is Like so RLVR is not actually about verifiable rewards like that has been failing since before the reasoning revolution like like like and that's the weird part about tasks right like back when oh fun history back when RHF was becoming a thing thing there were three different things that like are now called post training different efforts and instruction following was by far the like the the vaster child like people didn't like it they didn't want to take it into account it was annoying you know like I talked to the pre-training team.

104:37 I'm like, "Guys, this is the magic." And they're like, "We'd run so many model sweeps. You want us to wait for human evals to figure out which models to use?" And like everyone is like, you know, giving tons of like resources to like the codegen team, which like they did have some successes, but they were trying really hard to do RL on like unit tests, and it didn't work obviously, right? Like you needed reasoning for that. So re so just to be clear RVR is not purely about the reward. It's about like the shape of everything to and part of it is that reasoning is included in here like this latent variable that you're doing things and when you're doing things you're just letting the models do whatever they want in order to make them be as powerful as you can to answer the hardest problems.

105:22 And this whole pace the frontier discussion I think is like a very narrow focus because it assumes that everyone needs to do more RLVR, right? Which like I obviously don't think I need to do more RLVR on our models, you know? I think zero is the optimal amount for our shape, right? Like come on, you know? So it's really I think a bit of a slight of hand where they are saying that we actually want to keep doing the thing that looks dangerous because it does dangerous things you know like people say like oh maybe the sandboxing was a problem or whatever else. yeah, I mean obviously it is and they could have easily solved that, right? But they chose not to because the more things you let the models do in this do anything category, the more powerful it is, right? So like there I think there's some like disillusion of responsibility there on like things that by design or non-design they're trying to make is just an assumption. you know, we must do RLVR and not just we must do it. We must do more and more and more with giving the models like the power to do powerful, you know, do anything they want in the middle because that teaches them to be powerful outside of it and we don't want to limit those things well because it'll make it slightly less powerful on those things. so like if you assume all of that, they're like, "Oh yeah, we're heading into a dangerous world, guys." Like everyone is going to be doing this and this is the only way to make AI sick. So >> so basically it's like it's like these are all internally consistent, but actually starts from a premise that has alternatives. If you think about course I think there's like like on the bitterest lesson direction I think that there's very few people who've like made right tasks you know like new directions of AI that is or new new north stars that is rare again like I think 2.2 two times or something for LLMs itself like RHF and then RLCD RLVR is like a 0 2 in my opinion and I think that's generous but or.5 or like it could be one whole one I don't really care but I do think that people are thinking very closedmindedly about this type of thing and this the only people who are at fault here are the researchers because it's definitely not the populace you know like they just assume that open anthropic are just doing the best they can and they are not the experts who are aware of the true optionality available.

108:03 >> Yeah. And that's fair. And and like you're you're also doing your part in waking them up. >> Exactly. Well, I'm doing my best, but like my goal is not like convince labs that there's like other directions to to go down. My goal is have you know like spark hope in software engineers to start like actually automating things they've always wanted automated. I had this like article that I wrote that my team didn't let me write that didn't let me publish about like the the future I want of AI and like there's like a lot of like little things like remember do what I mean. Imagine if everything could do what I mean cuz like that that demo was do what I mean like like you like like say yeah don't do what I say do what I mean.

108:43 >> Yeah. And like we couldn't do what I mean yet because like computers are so basic and literal but that computer use one was just that. And I think that there's like levels of smoothness that will happen in the world that people just don't understand. And like the promise of like smarts all around are it it's it's I don't want to overpromise. I don't think it's going to happen right now, but like we are going to do whatever the we can to make that happen.

109:09 >> Yeah. any other things on the sort of general shape of post training? You know, you obviously you're been very intimately involved mid-training. Is that something that you do have comments on? I don't think we've ever talked about it >> mid training. I mean it's all a spectrum. >> Yeah. >> Right. Like am I >> This is a curriculum but like fancier. >> Yeah. I mean like it's it's it's like you know it's a cost-saving thing, you know, instead of like having to pre-train again. Like there's intriguing stuff. I actually think that >> like intelligence has a janisiqua at every single level and it's always super duper fascinating. like I'm a shape rotator so I don't like finding that but I love it when people find it and teach me about it. but and you know looking at the data this thing that our data team is so good at that I'm not it's I find it really really fascinating. I love actually thinking about like how capabilities are like put into the model like over like the short term, you know, like there's like the the the the really rapid alignment of fine-tuning and over the long term after seeing it over and over and over again like this stuff gets baked deeper and deeper and deeper and deeper into the model until it gets robust, you know, and that is like the the north star to surface and like the system one stuff is the stuff that ends up getting robust. So I find mid-rading to be like a fascinating thing. I'm a fan of all forms of training. I'm a fan of all forms of like surfacing new types of intelligence. I wouldn't do it all myself because it's expensive.

110:43 and I have said privately and also should I say this? Huh? You know like my philosophy is anything I I I should say in like private with like an investor I should say in public with a people because that is like my thing. Yes. So the thing I've said before is if you gave me a billion dollars I wouldn't pre-train. I still believe that to be true. It is a very expensive thing when if you are like like if you're an a engineer you can like slice and dice and do all sorts of stuff you know like Frankensteining is not the most elegant beautiful thing but it solves problems baby >> so anything except >> retraining yeah amazing I think one one direction that I do think that is interesting just like synthesizing all your all your commentary about these model things is like do we have a a a super model that has all these capabilities involved or do we break them out in f further right so like one way to put this is the openi was trending in the direction of the omni model right 40 was was one of those then for for for a brief period of time there's always like there was like kind of a main branch of this is the chat tune model and this is the coding tune model those are completely different things those are extremely different concepts I will like break that down a little bit so multimodality is a little bit different because sometimes the other modalities help sometimes they hurt like you know people are moving they seem to be moving away from speech which is different than audio because it seems to not generalize well to the other stuff this might get solved I'm a fan of all of this but these are like empirical real questions like scaling laws are not about just throw money at it and it gets good scaling laws are pragmatically how good is a thing you know like like there are worlds where like no matter what you scale it may not be good enough. So you know like computer use is not currently solved is my understanding like I'm hoping that we can be a like play a part in solving that but like there might be no amount of data we collect that will solve that we might need better methods or something else like that. So like we you need to be like really practical in all of this. Am I a fan of omnimodels? I'm a fan of all forms of intelligence, but I will go straight into one thing you talked about which is different from pre-training which is post-training cuz I hate fracturing intelligence. That is like the bad thing to me. And this whole like chat first reasoning mode is because it forces the intelligence to be fractured. Like when you're optimizing for chat, this tends to be like pure RLHF and it's quite intrinsic in RHF to do the stuff people like naturally complain about, right?

113:36 Oh, >> you're absolutely right. And you know, >> psychopancy, psychopanty, whatever word, how how to pronounce that overconfidence, hallucination, like even the kind of style that excels on Ella Marina. Bold, italicized emojis, you know, like it doesn't answer the question simply. It gives like a long write up and then it asks you a follow-up question so it feels more like a human talking to you. All of these things come because strings are super weird, you know? They are like weird ass things and you need to be miscalibrated.

114:07 You need to like mode drop. You need to be hyperconident in order to not go off the rails because the the reward model will punish you so hard when it happens because it's obvious. and then this like warps the probability space entirely and it interacts with that of the reasoning models, right? Because like it, you know, the the models are like these simple linear things that tend to cheat a bit. So I think that that's very different than exposing intelligence is is my guess and a lot of the art to intelligence is studying this subtlety that I think that at least when I was in open AI people were not really studying that because like they were just like chat chat chat chat just like people are with Jeff right now.

114:48 >> Yeah. you give me an ultimate, you know, you give me an objective, I will just go optimize for that, right? But but if you try and there and you know like the saying is like you could have like two objectives and you could just like optimize for both but then that is literally the act of fracturing right so yeah >> so I mean in some ways you have you are also factoring intelligence into system one system two but you just don't agree with the other people's fur fracturing >> which is it's a little different no if I could if I could add if I could defend the system two tasks number one like we don't toss out the system two tasks Right? Like you can try to make Jeff work on it and there actually is an intelligent answer for that which is unknown you know like like there like there is better and worse behavior in the system two tasks which should be like really low confidence lots of uncertainty maybe some heruristics can like move the needle here and there but we care about them too just to be clear I just think that that is not what the what is the intelligence is native to so we're not trying to fracture anything like that And all fracturing makes the model dumb. You know, like if people like get the model to say like it is OpenAI or Quen or you know like Claude or whatever else I don't really know what it says this these days. I am not going to put into the models that you are Jeb from Typesafe that fractures it right like it like I I don't want that like it represent what the internet thinks right like be correct that is what I want because that's how you get the smooth predictable intelligence. I I mean identity is a thing I guess that for a first party product yes but like for an API I don't think so you know like I I don't like people don't want if they're making a chatbot with you know chat GPT they don't want to say it's chat GBT they want to say it's like chipot lei or whatever right >> well you know so the way that you also have to make up for it is you have the skill right the the the Jeff skill which which is for coding agents to work with Jeff >> okay a couple of closing questions because I I I do to get you out.

116:50 one is like just reflecting on your two-year journey. This is roughly two years, >> 2 point something >> with the company. I think that this is like more like a four-year journey, but yeah, actually like I was thinking remembering that like you had this like hero run around Thanksgiving. You were like you were canceling everything because you were like guys like everyone's on holiday. I'm going to take all the opening GPUs and go do this thing.

117:11 >> Yeah, >> that was a good time. >> And that was like the the pre-type safe moment, right? I that might have been Was that when the coup was happening? I don't really know. >> Yes, actually. >> Yeah. Yeah. Yeah. That sounds right. Yeah, I remember. Oh my god. I don't want to I I'm not I I don't think I have the time to spill the tea about the coup right now, but that wasn't really annoying.

117:33 >> the the coup was annoying or the the run was annoying. >> The the coup was annoying. >> The Okay. Yeah. Yeah. >> It I will >> safety took over the company. >> Yeah. Anyway, >> maybe next time we chat I'll I'll dump tea about a tea about the coup. yeah, it actually this problem was one that like was in my mind since before chat GPT even launched. I was like, "Holy the chat GPT team is cooking. They are doing the right task.

118:02 They are doing the thing that AI researchers are bad at but successful product people are good at, which is giving a lot of about the experience." you know, it's it's it's very rare. Like there's very few people like that at OpenAI. and those guys were cooking on it really really well. >> And to be clear, this is the whole journey from GPT 3 to 3.5, which included AI dungeon, which you've talked about as like Yeah. Well, that's that's an example of a use case that we never predicted.

118:28 >> Yes, exactly. Well, oh yeah, that is also I had fought very very hard to deploy instruct GPT. like actually the early versions of it were even trained with like an algorithm we didn't publish that I made myself because it was too slow to clean the the PO data and I was like it. This is so good. We need to get in the hands of users and like basically immediately it took 50% of the market share of LLMs at the time and but and we thought I I made I went through great effort to make sure everything in our launch video is true. you know, we I truly was thinking like is this AGI because it's super human at instruction in instruction out. You obviously it's not but like everyone I think should have an answer to why that was not AGI because it looks very smart. And my answer to that ended up you know like ended up only being used for copywriting. You know Jasper AI copy AI like writing like you know what is now called slop on web pages. and we were worried we made the internet a worse place right? And I went back to the drawing board and I was like what's missing? We are smart clearly. Something is missing from it like creating value.

119:36 What is it? Like I actually was doing more philosophy at the time of like you know like what is going on? And the answer was oh machines. You know the question I asked myself is like let's work backwards from an AI based economic revolution. When that happens what will be what'll be calling the AI if AI is an API? Will it be humans or it'll be code? And I figured it was many nines of code and but like all the optimization was going into the humans part and then it clicked for me. I'm like holy this is the northstar. I think like I wrote a document. I was like talking to Sam about this. Sam was like this is so good. You should go work on it.

120:15 And we're like yeah yeah Sam I have a job. You know like you know I was working on >> Sam just told you to do it. Do go do it. But like my and my guess at the time is like this is super obvious. Like it's so unbelievably obvious Anthropic must be working on this already, you know, and like we're already cooked and like actually open does better at like catching up than it does at like actually innovating. So like chatd was a copy of Claude, right? like they had an internal thing. They just didn't ship it.

120:42 >> Yeah. Cloud and Slack. But but you know reasoning I I would say firstish. >> Yeah. But debatable how good of a product that is. Yeah. great research though. Super great research. I'm just not sure if people had that product need. and you know, Claude did the coding agent stuff too. So Sam says that and you know I just go back to my job for a while. Eventually like you know the instruction following team just says we won. We we've solved instruction following. We don't need to do stuff anymore. I'm like trying to think about what I do next. I'm like you maybe I'll just start start playing around with this. I you know do more philosophy and design and thinking. I thought it would end up taking a week when I started training models. It ended up taking many years. At some point I was like holy you know there's signs of life here. This it obviously didn't work right otherwise we would have deployed it. But like I want to explore what it would be like research- wise to go all in on this, you know, like I want to really see like like what it would be like if you went like absolutely insanely all in in this direction. And because of what I said, you know, like if an AI winter happened, would I how would I feel? I would consider myself personally responsible.

121:56 I talked to other companies at the time and I was like, "Hey, I want to start a lab on this direction and you know like there was interest and I just talked to them like how fast what would be faster this or startup and they're like startup and I'm like it man we ball I guess we're doing some crazy and >> and you call Eric and Sasha and >> Yeah. Well, I I call Eric first. with Sasha I actually didn't try to recruit her. I I try to be good and I was just like, hey, am I crazy? Is something missing here? You know, isn't there like like like am I too much in the OpenAI bubble that I didn't realize there must be a solution to this? And then Sasha was like, I'm in. And I'm like, Sasha, you're working at a startup. And she's like, I'm folding it right now. And I'm like, do you want to think about that?

122:44 She's like, oh yeah, good point. Let me think about it. And then she joined. And then, you know, within 2 weeks, we had funding within we had like people move into my apartment. It was the worst cuz I'm a neat freak. and we just kept on cooking and eventually we got the research that that that showed the signs of life. You know, it was it was a crazy time. So the the the the question is that was all long context and now the question is someone like you is in the frontier lab right now who is frustrated not getting the the funding or the the resources whatever the attention what's your advice to them you know do should they do what you did should they do it that's a fascinating question man how do I do this without burning bridges I my sense is that most unless there's some level of economics I don't really understand. I think most Neolabs are crap. I I don't want to see myself with that as peers. like I I don't really understand what's going on there. Like is it because like number one I don't really value researchers.

123:59 I value people who like like look at my bitterest lesson, right? I want Well, not just that we need researchers, but we need them to give a lot of about the right task and that's the important thing, >> right? So, it's actually like the it's it's kind of backwards when people value pure research pedigree because that generally doesn't create value. so it like number one I I believe in northstar tasks and doing cool really useful stuff. number two, because I don't value researchers, I don't I don't recommend going the well, it clearly is profitable for someone or it might be in this environment. So like from a purely pragmatic perspective, I don't see creating Neolabs as one something that creates value. It seems to destroy value because like they are like redoing work from scratch with like low probability of actually moving the frontier. And as far as I've talked to most LEO labs, they don't really have a direction. They tend to want money to play around with their experiments.

125:02 if they have a direction, I'm super in favor of it to be clear. So my advice for someone is it really depends on why you're doing it. You know, if you are a researcher who wants to play around with research, probably the labs are the best place to do that. TBH like there might be other places. I I don't really keep track of that politics, but I would just recommend not being that way personally, you know, like I think it's better for the world with people being driven to solve real problems. And those problems may be exploratory. That's fine. But like ideally have principles that you stand behind, but if you think that you want to do the right task, like abso lutely, like please do like please break this like uni mind, you know, unimodal like mind. Exactly. Like like you know like again this pacing the frontier is coming from like this one view of AI that looks like you know AI super genius that is incredibly jagged and that is >> solvable. It it's solvable and it's weird and it's like not matching reality and it's like it's tragic, right?

126:14 Like like I I I think like all of these like that like really unearthing technology I think is like just good. >> Yeah. For what it's worth, you know, again, I'm trying to rep accurately represent the position of the, anthropic open AI folks. I was talking to, SpaceX as well, by the way, is that, it is this is a political thing much more so than a pure X-T thing. >> Yep. >> so yeah, political positioning >> and that and that's beyond my that's well beyond.

126:40 >> Once they once they told me that, I was like, I get it. this is about the 2028 election. >> Oh, no. Oh, I wish I didn't hear that. That's such a bad vibe. >> No, no, no. This is not the whole company. This is just that that room's >> discussion. No, no. That that makes sense. That that that makes me lose faith in humanity a bit. But maybe I'm just a naive technologist. >> it's really starting to matter who's who's in charge of the governments that that will help to regulate these things as they emerge. And like as a lab, you should probably think that through.

127:12 >> No, no, I I totally agree with that to be clear. Like I I think being opinionated on that matters a lot. I personally am afraid of trying to mislead people because I think that bites people in the ass a lot, you know, I think that like people trying to be overconfident like I I I I obviously I'm not actually going to talk about politics. I think what happened in CO is like people leaned too much in like appeals to authority and being overconfident to try to get people to behave in certain ways and like obviously our response was extremely suboptimal. And that had like like ripples of downstream ramifications that are now I think extremely bad for the world. Like maybe I'm naive. I think that misleading people even for the greater good or what they think is the greater good is just it's just I'm not a fan. I rather not >> for what I it's not I don't think it's misleading. It is just like this is why now >> why yeah like you know >> I I think like how come that is why now that is a little bit misleading about like the risks versus like the the objective.

128:27 there's like some level of like sneakiness latent in it that is worth calling out and I think owning up to. Well, obviously they want to if if they want to manipulate then they shouldn't own up to that. That seems like a bad strategy. But like that to me is just just sad for the world. >> Yeah. >> Hopefully I'm never Yeah. hopefully like we are never involved in anything like that. It might be inevitable as we get big. but I want to I want to stay like pure technologist to my roots as much as I can.

128:58 >> I mean, Jeff for president, why not? I can, you know, I trust Jeff's decisions over my own. okay. So, so less posting. more about you're just crushing my hopes about America and the world right now. >> Oh my lord. >> Yeah. I mean, like there's I think I watched too much TV about like conspiracies to take over the presidency. the you have chosen your northstar, you have chosen reliability and then programmable and composable AI and cheap.

129:29 >> What is a second or third one that you want to throw as a bone to someone else that you're not that you want someone else to work on that you're not going to work on? >> Oo, >> like just basically give people tasks. >> Give people tasks. >> Yeah. Like like your task, >> there's so many I want. >> You have picked your tasks, right? You know what I mean? >> What? Wait, that's such a good question.

129:46 Holy crap. Oh man, I'm so excited by that. cuz you're you're going to be you're be the next like 50 years you're going to be busy doing your thing. >> Hell yeah. Okay. So, let me give like a fun one and a not fun and like maybe a valuable one that's also fun. my my fun one is I think games could be so freaking cool if they were intelligent. Like when I see people play around with like like Alli's Doom demo where like you can like get NPCs to control stuff like you know like that was just really like the like a proof of concept. I think some really cool stuff could be made. It looks really really cool, you know, like like I'm a big Stardew Valley fan, you know, and like it's it's really static and it's still compelling.

130:30 Like I feel like there's a lot of cool story that could happen. You don't need to call like Jeb in the game loop. It's probably too expensive for that. But even like simple like state machines for NPCs, I think you could make like such a compelling world. Oh man. and man, a little sad that I can't work on these types of things. my my life path is a little bit set right now and I'm >> Yeah, but you can call someone else to work on it and then you can like feedback on it.

130:54 >> the thing that I would really really like to explore is like coding agents free from the tyranny of the KV cache. Like like it might not be as good as true coding agents are, but I think there's just so many weird things to think about. That's why I wrote the article KV cache rules everything around me. believe it or not, I don't think anyone has used the phrase on the internet cash rules everything around me. C A C E. when I when I Googled it. so like I wrote this cuz I wanted to tell people about like this is how coding agents agents work and how the KV cache works and everything. And I think oh yeah like it explains a lot of stuff like why routing is really hard why sub agents don't seem to work like why compaction is such a hard problem and I I'm going to try to release a document my team might veto me because believe it or not I'm not in charge you know but I wish but I want to release a document like here are my thoughts Please play with it and please figure out all the ways that we can do things with coding agents like once you're freed from that you know that that KV cache tyranny >> which is it locks you in >> well not it lock it it locks you in into one model right and in order to do it efficiently you need to like keep on appending to it so now you're not doing best software practices like state management abstraction decomposition why can't you give an easier task you why why can't you give a sub agent easier task because of the state that you're passing around. Oh I touched this because of the state you're passing around you would need intelligence that is way cheaper than the intelligence using to read this in order to pass this state around. Why can't you be smart about it? Right? And I think there's like tons of really cool fun research to be had there on like different programming patterns. You know kind of like how people are playing around like with like recursive language models.

132:59 Like I feel like there's like just lots of cool stuff in here. When you think about like oh I want to explicitly label the state of everything or imagine you have like a subtask like coding agents I I think it's fair to say they work on subtasks at a time as from a decomposition perspective. Why do you need to pass all of that state back into the parent task? Why couldn't you do smart things about it? And and also if you had a hierarchy of labeled subtasks, why can't you do a search through that subtask tree for the relevant context when you need it in, right? And then you know another thing that you can do, oh man, I forgot to write something about this. I have like some cooks in here that are really really cool. hope to publish it. I'm down to jam about it, but like it's going to be a long document. And and like if that becomes the case where context becomes cheap, like why can't you do cool patterns like looking at your historical context very cheaply? Isn't it kind of weird that you start from scratch every time and you need to solve a problem called continuous learning? That's a that's actually like a memory management problem because you don't have a smart way of looking up the memory, right? But what if you could what if you could do that all the time? Or what if when you have parallel sub aents, they can like read each other's states because you have all of that in like your computer memory and you can be smart about what's reading and writing at the same time and your coding agent swarm or whatever has like locks around things and can coordinate intelligently not with like basic ass locks like what are you doing what am I doing you know Jev who should write first blah blah blah and like I feel like the future there is nuts >> Jeff to solve locks >> I mean it could be so cool for like multiple agents working together or like if you think about state like agent forms.

134:42 >> Yeah. And and you know some things for example are read only processes. You know some people like getting like summaries of what the agents are doing. Why can't they share state easily because like a readonly agent needs to like you know read parts of the context and figure out what's relevant to say like what's actually being written because exploration is not super important or here's the tree of subtasks. I feel like there's so many different fun things that could be done if like a really smart person like dedicated like a whole lot of time to rethink like the the coding agent experience and that would be super duper sick.

135:15 >> Yeah, >> man. That would be my dream. >> I I would point you towards Prime Agent if you haven't looked at it. so this works together with the RLM work. We just talked to Alex who is a buddy of Allen's in the chair before you. Cool. and like yeah it is being worked on but it's not super popular yet and if like >> well yeah but the hope yeah I would want everyone to like just play around with with like weird things. I have no guarantees that it'll work but it seems really really interesting from like a technical perspective.

135:45 >> So yeah that that seems cool like like once we figure out how to give credits out I would love to like give credits out to people like this. Yeah, you you will be in a position to fund research for sure. no, anyway, congrats on all your success. you you've like come such a long way since I first met you like and and and the whole team >> the same person as well. >> Yeah. Yeah. I I think but I think like you are energized in a way that I have never seen you before because you found your mission, you know.

136:09 >> That's true. That's definitely true. >> and you articulating your mission because you you you for many years you complained about the problems but you didn't have a solution yet, right? and you're like you had you had the rough shape and then you had to do it put in the work. >> I will say that that is partially because I I describe myself as as zeroth% entrepreneurial. I I I don't like startups. I never wanted to be a CEO in my life. I can't imagine anyone doing this twice. It seems horrible.

136:41 Honestly, doing it once is pretty bad. when we first were fundraising, an investor asked me like, "Which CEOs do you look up to?" And I was like, "Ew, why would I look up to those people?" no offense to anyone, you know? I'm trying to be like I'm trying to be genuine good and I've met like a lot of really good people, but like the famous ones have like a lot of like skeletons in their closet, it seems. And I think I just really did feel disempowered when I was at OpenAI, you know, like I felt yeah, like like like it's a little bit easier to be truthful now because like I have at least some proof that the direction has legs. Like I just felt like in the the the insane house where everyone is just like chat GPT. Yeah.

137:21 Like where do we put chat GPT and everything? How do we make chat GPT good for like you know developers and stuff? And I'm like what what are you talking about? Like the function calling interface is insane. Why would you deploy this? like this is this is so anti-developer. >> Sort of a hacky way on top of hacks on top of hacks. Well, not just that. Like the thing I often said was if there was like a this is also probably t I don't have time for right now but I always used to say like I want to be removed from any project involving like function calling if you did not get a legit bias for each function like so very very simple ask in my part >> which is something like a confidence but not calibrated >> or a probability for it right like we need to give users the ability to control like you know like let's say the actions are refuse or allow yeah Disney needs to set a different refusal threshold in AI dungeon the only way to control that with function calling right now is to say like pretty please you know that's nuts that's a nuts interface for developers and like people have been like dealing with this for years now right like they still have that with skills like you know like the existing coding agents are like highly overfit to their existing harness because they're jagged They don't tend to use like external like tools and MCPs super well because of overfitting of course and like why can't like big companies allow for like slight nudges to be like call this more it's really useful right and and like the the the the solution is begging in a system message that's nuts. but okay I I think I I think I get you and like man it is so exciting to to talk about all this stuff. It's it's really cool to get you on the podcast. Yeah, you're going to go do amazing things, man. Like I'm excited for your next big launches.

139:13 >> Oh, hell yeah. Just you wait. >> Just you waited sooner than you think. >> Infra people, I assume. Marketer >> depends. If you ask me, I feel like I'm a pretty good founding marketer. But if you ask anyone on my team, they say, "Shut the up, Yogo. You need to do CEO stuff." So yes, founding market. >> It's not just about spice. It's like I think you're very spice oriented which like you like that's your unique talent but sometimes you just need to say >> I know I know marketing things. Yes. I nothing teaches you delegation like having a title wave of stuff to do.

139:47 hiring data people or we call them model capabilities like but they are data people both like like data is kind of a slur in the industry and like I want to make sure >> so we're very pro data here. >> Yeah. But I want them to be the highest status of like, you know, the people actually working on the model. That actually sounds a little weird. I want everyone to have equal status, but like I want to even that out and I want to know that that's really valuable.

140:09 >> At least I'm more equal than others. >> Well, I don't like weird hierarchies. And I I think one of the things I'm most proud about in the company is that they don't respect me that much or they don't show that. They just troll me and like joke with me and they treat me poorly sometimes and all of that. And I think that that's a good sign of a culture. We're hiring like platform people, like people to like build out Jev everywhere.

140:34 Like we are so much more sensitive to location because speed of light is more of a bottleneck, right? Like I'm so sad for the European users that we're only like three times as fast instead of like a 100 times as fast because like we don't have servers there right now. And that's insane, right? >> It's okay. Life in Europe goes a bit slower as well. It's okay. >> Wow. I can't believe you you said it, not me. or or everywhere, you know, like if if intelligence per second is a metric that matters, like we we'll want this all over the place. Like we care about like if they're a developer building on top of us, I care a lot about you. and we are hiring for people to keep building more like not just like the goal is not to just be like Jev as a company. The goal is to like ship more shapes of intelligence beyond that. So we are hiring people to like build those things too, you know, like we want to not just be like yeah like the one trick pony of like the simple model, but like I think that there's going to be like an AWS of like intelligence, you know, >> which is going to be you by the way, right? Yes.

141:38 >> I mean like that's a direction I want to go down. It would be arrogant to say it will be me. Like we like I'm going to do anything I can to make sure that happens. Like I think that that's going to be so so cool, you know? Like we are playing with like system one intelligence right now. Imagine the layers, you know, like this is like the TCP of it. >> Yeah. seven more layers to go and who who knows what else. I've also pitched temporal by the way. I don't know. We need to talk about temporal as layer eight out of the seven layers.

142:10 but anyway, >> we can talk forever. You got to get back to work or sleep. thank you for coming. >> Oh boy. Yeah. Cool. You're most welcome. It was a pleasure, man. >> So excited. So excited. My first time.

Summary

The discussion revolves around the launch of Jev, a new AI model from Type Safe, and the broader implications of AI development. The conversation highlights the challenges of automating basic tasks despite AI's advanced capabilities, the importance of community engagement, and the need for developers to be empowered in their use of AI technologies. The host and guest reflect on the evolution of AI, the significance of model reliability, and the potential for AI to drive economic change.

- Jev is a new AI model designed to optimize for intelligence per dollar, aiming to integrate AI more deeply into software development.
- The conversation emphasizes the disparity between AI's advanced capabilities and its inability to automate basic tasks, questioning the economic structures that hinder this progress.
- Community engagement and developer feedback are prioritized, with the guest expressing gratitude towards developers for their support and insights.
- The discussion critiques the current AI landscape, particularly the focus on safety alignment and the limitations of existing models in handling complex tasks.
- The guest shares a vision for future AI developments, including the potential for coding agents to operate more effectively without the constraints of traditional models.
- There is a call for a shift in how AI is perceived and utilized, advocating for a more practical approach that prioritizes real-world applications over theoretical constructs.
- The conversation touches on the importance of data and model capabilities, suggesting that future advancements will rely heavily on innovative data management and processing techniques.
- The guest expresses a desire to create an ecosystem where AI can seamlessly integrate into various applications, ultimately enhancing productivity and creativity.

Questions Answered

Why can AI solve complex problems but struggle with basic tasks?

The discussion highlights the paradox of AI's capabilities, where it can tackle complex mathematical problems yet fails to automate simpler, economically valuable tasks. This discrepancy raises questions about the current state of AI technology and its integration into practical applications.

What are the ethical considerations for AI API usage?

The conversation addresses the ethical implications of AI usage, particularly in contexts that could lead to harm, such as warfare. While there is a desire to prevent misuse, the speakers argue against imposing restrictions at the technological level, as it could compromise the AI's intelligence.

What does the future hold for AI and its impact on technology?

The speakers express optimism about the future of AI, emphasizing the need for continued exploration and innovation beyond current models. They suggest that the advancements in AI will lead to significant changes in technology and the economy.

What challenges does the AI industry face in development?

The discussion reveals concerns about GPU constraints and the need for efficient AI models. The speakers highlight the importance of making AI accessible for experimentation and innovation, likening it to a gold rush in technology.

How can AI codebases be improved?

The speakers discuss the potential for AI codebases to become more efficient by breaking down problems into simpler components. This approach allows for better evaluation and verification of AI outputs.

How is AI contributing to economic change?

The conversation touches on the economic implications of AI, suggesting that while it may not lead to mass unemployment, it will cause significant shifts in the economy. The speakers advocate for AI to enhance everyday life rather than dominate it.

What are the practical use cases for AI?

The speakers discuss the importance of reliable AI applications and the excitement surrounding new use cases. They emphasize the need for thorough testing and understanding of AI capabilities before widespread adoption.

What is the philosophy behind AI model development?

The speakers share their philosophy regarding AI model training, advocating for a pragmatic approach that prioritizes problem-solving over traditional pre-training methods. They discuss the balance between developing comprehensive models and specialized ones.

What are the political considerations surrounding AI?

The discussion highlights the political dimensions of AI, particularly in relation to governance and regulation. The speakers express concerns about the influence of political agendas on AI development and its societal impact.

© transcribe · For agents Built with care and craft by Gokul Rajaram