Section Insights
Understanding Agentic Engineering
What is the importance of being ambitious in agentic engineering?
The speaker emphasizes the need to recognize unknown unknowns in agentic engineering and the importance of being ambitious in learning and prompting to create valuable work.
- Agentic engineering involves more than just simple prompts and loops.
- Recognizing unknown unknowns is crucial for progress.
- Ambition in learning can lead to better outcomes in AI development.
The Non-linear Growth of AI Models
How do AI models improve over time?
AI models do not improve in a linear fashion; their growth is often unexpected and can emerge in surprising ways through experimentation and exploration.
- AI models can exhibit spiky improvements rather than consistent growth.
- Understanding model capabilities requires experimentation.
- Growth in AI can come from unexpected methods, such as tool calling.
Discovering AI Capabilities
How can we uncover the capabilities of AI models like Claude?
By actively engaging with the model and asking specific questions, users can discover unexpected capabilities, such as video editing, that they might not initially consider.
- Engagement and specific prompting can reveal hidden capabilities of AI.
- AI can perform complex tasks like video editing without traditional tools.
- Users often underestimate what AI models can do.
Addressing Unknown Unknowns in AI Development
What strategies can be used to address unknown unknowns in AI?
To tackle unknown unknowns, one can ask detailed questions and seek visualizations to better understand complex topics, such as color grading, which can lead to deeper insights.
- Asking detailed questions can help clarify complex subjects.
- Visualizations can enhance understanding of abstract concepts.
- Engaging with AI iteratively can lead to better comprehension.
Evolving Capabilities of AI Tools
How have AI tools evolved in their functionality?
AI tools have progressed from basic question-asking capabilities to more complex tasks, such as generating HTML reports, showcasing their increasing sophistication and adaptability.
- AI tools have significantly evolved in their capabilities over time.
- Switching formats can unlock new functionalities in AI models.
- Understanding the evolution of AI tools is key to leveraging their full potential.
Transcript
0:00 I I think it's just more like the highle idea of like you probably have a lot of unknown unknowns, right? And you're probably not being ambitious enough. I think we sometimes oversimplify agentic engineering where we're like it's just prompts, it's just loops or whatever. And I think it's more like no, we need to know what we're doing. We need to like learn more and then like be better at prompting and and make sure we're creating like valuable work.
0:26 >> >> All right. Okay. Cool. Hey guys, my name is Stark. Is is the mic working fine? Yeah, it's good. cool. Okay. Yeah. thank you. Thank you. Yeah. yeah, I work on the cloud code team and yeah, I I I've been working on it for about a year now. Before this, I was a a YC founder and ran that company for about five years. decided to get into AI. Actually, Eric is here from Goodfire. Eric was like my first sort of like AI gig, you know. So, like I did some interpretability work with with with Goodfire. eventually ended up on the cloud code team. and yeah, so I Okay, I I thought there were like going to be seven people at this event. And so I I was sort of like, you know, you have to sort of like shape your content on on on the audience. I I'm super excited to talk about this. I I have a bunch of ideas I want to talk about. and I have a Yeah, I've had giving an AI engineer keynote next week, but this is sort of like mid formation of these ideas. So there are like a few different like things I'll I'll jump between. So like please like, you know, bear with me a little bit. I I think there's like a good thread of thought, but you know, like maybe everything won't be like the decks won't be color aligned or something, you know. So, and I'm not sure if this is the title I'll go with.
1:49 I'm I'm not sure, but okay, the the primary thing I want to talk about is like we need to be better, you know? I mean, like I I think like you know, people tag me on Twitter sometimes, they're like, "What are you spending all the tokens on?" You know, like what what what's happening? Where's the like where's the growth? And I'm like, you're right. You know what I mean? Like like we like I think the goal for AI is to like really meaningfully improve, you know, like how like humanity, you know, like like our GDP, right? And I think that we haven't shown this yet.
2:22 and so how why right I I think the thing that we all obviously see is like software is becoming super cheap, right? Knowledge work is becoming cheap, right? and this is like you know like obvious to us but I think maybe less obvious is that like generating value is still really hard you know and like building startups as I'm sure many of you are doing here is extremely hard and bu software and knowledge work are parts of this right but it's like not all of it right and we haven't honestly earned like our valuations we haven't earned the like you know the amount of money we're putting into AI yet I think you know because there's still more value to create, right? And so, why, you know, and I I think like the the thing I come back to that I've really appreciated at Anthropic is like we have this value that's like we don't negotiate against ourselves, you know, and so I remember being a CEO of a company and being like, okay, what are our priorities? Let's write down the priorities and then let's figure out how they trade off against each other.
3:26 and it was very reasonable, but like, what if you were less reasonable, you know? I mean, like what if you ask reality to kind of show you what the trade-offs are? And I feel like what I really appreciate at Anthropic is we're like, let's just do the thing, you know, let's just do all of it. Force us to be shown what the the trade-offs are, right? And I think with AI and Claude more than more than anything, I think it's like, you know, what are the trade-offs really? It's harder to tell, right? Like is good, fast, and cheap still a trade-off? maybe not right and so I think that this like the qu like you know like how do we free ourselves from the mental models of like you know what our trade-offs are what our ambitions are things like that so that we can ultimately deliver the promise right of of AI so this is something that I've been thinking a lot about and and you know I'm kind of like okay why why you know why is this so hard I think one of the reasons is that like LMS are just really weird and they're like a new thing and we need to like figure this out, right? So, okay, whatever. yeah, I think it's like I think this is like roughly the art of like human agent interaction, you know, and like it's a very I I think a lot of the, you know, tech companies in like the 2010s were built off human computer interaction, right? Like UIs, like UX, like incredible like patterns and like now we have to have develop this like new like technique, right? Human agent interaction. and I think like talking to LLMs is like is a new skill, right? It's like prompting or public, sorry, it's like public speaking or writing or any of these things. if you're good at it, you can drive like a ton more value than, you know, someone who's not good at it, right? And and so it's something that I think will continue to be like high leverage for a very long time. and I think it's important for us to I I think acknowledge that, right? Like I don't think the end goal is like you know, you just put a sentence into cloud and it like does the thing for you. I think there's a lot of like depth here. and it's because the models are like grown not designed right so you we like are cultivating the data the RL environments the like you know etc etc like all of the things that like go into pre-post mid training but we don't know what the end outcome will be until we get the model right like it's like you don't just decide like hey this model is going to be you know 95% on bench you grow it right and we you grow with a lot of care. but they're organic things and like we don't exactly know what they will be good at until like we really try them, right? we have some guesses. but these things emerge like in spiky ways. And so I I say that like models don't get smarter in a straight line, right? They get smarter in unexpected ways. I think one of the like best examples is that you know if you were to go back to like maybe Sonic 3.5 and you look at cursor and you're like okay how does the model like solve coding you're like okay obviously the context window get really really large like a 100 million token context window and then like we just fit the entire codebase in and just solves it right like that's how you think the model would get faster but but it didn't happen like that right it like got better by doing tool calling and by bash and GP and and like how do you figure that I don't know. You just have to you have to do it, right? So, I think they're smarter than they think and we're hobbling clawed is what I think a lot about is that like part of the reason that you know we haven't achieved as much growth as we could is that we like you know wow we are hobbling clawed. so I'll give a case study here that that I like. Pokemon ending in awe. There was this tweet about like, you know, this font may be too small, but basically like there was like an AI hate tweet going on where I was like, I can't believe chat GPD can't answer Pokemon ending with AW. You know, Cloud could answer it, but like we let's not get into that.
7:22 but but well, it's why, right? So, there are over a thousand Pokemon. exactly two of them end in awe. It's Crocana and Dreadnaugh. I didn't know that. I'm a big Pokemon fan, but like you know most Pokemon people would not know that a priority, right? and so chatt found one. and so like how would you solve this? Like like let's say we obviously think the models are smart enough. Why why couldn't it? And how would it solve it? One idea is like it's all in the weights. You just remember the weights like you know like obviously it knows every Pokemon. It should just think fast enough or like it should just remember or it should like think you know and just think out all the words. or it searches the web, right? And like all of these things like it it will not give you the answer, right? This is kind of empirical like is there a version of like the architecture that does just remember maybe you know but like empirically this does not happen but the models are obviously smart enough. So what works is like you ask cloud code you ask any like coding any agent with a code exec tool right and cloud code will like get the list of all Pokemon and GP for once that ended aw right and so this is like one line it'll just do it and like if you were you know like if if you're an average user you just don't you're like this you ask you know a model why what the Pokemon ending in aw you are and it doesn't and like you you just stop there, right? Like you just don't have the ability or you don't have the skill of understanding claude well enough to be like, oh, like how do I prompt it better? You know, how do I like give it the tools needed? Right? unhobling it. So, but you can go further. You can generate a web app that like you know searches any like combination of reg x for Pokemon, right? And like like there's so much more abundance than than you than you think or like than the problem implies.
9:21 so yeah we call this capability overhang right like we call like the idea that like what the models can do and like what we like you know are utilizing them for are mismatched right and I think definitely like you know open 4.8 data has like incredible capability overhang to me like fable like you know just I feel bad about it right so I I think like yeah the but but why is it hard I I think like one of the things to do is like let's try and put ourselves in Claude's shoes right and so the usual framing is like it's just the model in the harness like you just you figure out like you know what the tools are and you put it in but like let's say you're claude like it's more of a complicated relationship between like the model, the harness, the world and and you, right? So, yeah, it's like putting ourselves in its shoes. So, let's say that you wanted to like explain how the odds module works, right? This is like a very this is maybe even a good prompt that someone might put in, right? but in order to do this, like it needs to know who you are, right? Like do you know anything about the codebase? Are you technical or not?
10:29 You know, like how technical are you? what level of depth do you want? the world, right? like how big is this codebase? This actually has a big implication, right? Like if you're doing this in like a large legacy codebase, you need to spend a lot of compute. If you're doing it in a small codebase, maybe you have no sub agents, you just you just do it, right? And so this is something that cloud needs to figure out too.
10:52 and then like what other context can it get, right? So can it look in git or slack or like you know like is there like more nuance around around the O module? Why is the user asking me this? Right? And like if like your boss came up to you to like, hey, explain me what the O module is. You might be like, well, what for what, you know, I mean like how how do I get more context so I can answer this question in the way you want, right? but claude doesn't really like get the for what, you know?
11:20 and so yeah, and of course like the harness can help with some of this, right? Like it can do memory, it can like learn a little bit more about you. but there's still like a lot here that like Claude, you know, like we have some work to do. and and so like, you know, one example of making this prompt more explicit is like, hey, like I'm an experienced TypeScript engineer. I have zero familiarity with this module, but like it's kind of a high compute problem, so use sub aents, right? So, this is one way that you might get like a little bit more precise.
11:52 now, okay, let's see. I think like I kind of want to switch tracks a little bit here right now and I want to talk about sorry this is one of those things where I haven't figured exactly okay cool yeah yeah so another thing like unhobbling claude putting thing putting ourselves in claude shoes what are like you know how do we discover what claude can do right I think one of my I I think there's a lot of room for surprise here. And so one of the things I've been recently well surprised about is video editing, right? So Claude like just edits videos that we now, you know, like I shoot them with the like a video agency and then it end to end I don't use a video editor. I just use cloud code and it generates videos that you know look like look like this, right? So it's you know showing me like it's generating the UI as well. It's generated this across many different cuts and so it's decided like you know which are the best cuts of the data. it's like cut out ums and things like that, right? and it's, you know, done some pretty like impressive like UI, you know, like it's done these overlays and things like that. And yeah, I'd say like if you were ask people, you know, hey, can Claude edit videos like you know, they they they wouldn't think it could, right? and like how do you even like like what does the process or if you were to ask Claude to edit a video, it probably wouldn't be able to answer you in this way, right? because it it just doesn't you know know how to unhobble itself. So yeah how does that look like right and how do you get to this?
13:38 I will so the high level is that it's code generation right and so this is like a representation of the folder I got here and some of the previews. Yeah. I mean like the the idea I want to leave you with is like there's a bunch of clips here. Oh, maybe I have it here. Okay. Yeah. there's a bunch of clips that I'm given. and this is essentially the raw material of the like video, you know, and then what I ask cloud to do is I I give it a transcript as well. This is the transcript I was working from. so it has an idea of like, you know, what I'm trying to say.
14:17 and I ask it to transcribe it, right? And so it does, let's see, it does a bunch of transcriptions. there's so much stuff here, but this is like essentially what knowledge work is increasingly becoming. It's like just a folder of code and scripts and data, you know, that is like like controlled by an agent, right? And so, yeah, roughly what it does is, and the funny thing is I I asked it to make this, artifact as well to show you guys to sort of simplify the the process.
14:57 but yeah, it makes a bunch of, transcriptions of the of the of each video. It decides then which clips to do, which map up to my transcript the best, and then starts editing them, clipping them together, and then making UI to go along with it as well. So, it's generated a bunch of UI here in using React that will then compile compile together into like the end video, right? And it's also deciding which UI elements to make given the like Figma and you know, design system from our team.
15:35 more than that. So, I did all of this the first pass. This was the first video I did. and this was the second. And I think there's like a big difference here is like the color. So, like one of the things I had no idea about is color grading, right? Color grading is I still I know more about it now. but like you know when you get a video like turns out they give you in in this like raw format which is like sort of overexposed and and there's actually a lot of art in like oh how do you bring out the color of a video right? and this was something where this was a place where I realized I had a lot of unknown unknowns and okay let me come back now I'm like sort of deciding how how I want to like so yeah what does it mean to be good at agentic engineering right like I I think like how do you like sort of how do you solve these problems like how do you like yeah what's the skill like and I think it comes down to a lot of like unknown known, right? And like with like there are known known like what do you want? There are known unknowns, what you haven't figured out yet. Unknown unknowns, things like that are obvious that you don't know yet. And sorry, unknown knowns that are things that are obvious but you haven't you only recognize it when you see it. And then unknown unknowns. And so in this case, this was me being like, wait, I have so many unknowns about color grading, right? This is one of those things where in order for me to be an better agentic engineer, I know a lot about how ffmpeg works, how video transcription works, how reotion works.
17:07 I know all these things and that's what let me get far enough to edit this video. but I didn't know enough about unknown unknowns, right? or about color grading in particular was like a big unknown unknown for me. And so the question is like how do you fix that, right? And I think this is like a big problem overall. Like we if our goal is to ship better things faster to like, you know, drive GDP, to make better products, we have to get really good at grounding down our unknown unknowns.
17:37 and so what I did was I ended up asking Claude to tell me about color grading. And there is a version of this where like sometimes you just ask Claude and then you like sort of glaze over the like the report, you know? This happens a lot. I I think it's kind of like education porn sort of. You're like, "Oh, like yeah, this is, you know, it looks cool." But I I I it generated this and I honestly had no idea really what color grading was from this. I couldn't prompt it better. I couldn't like figure out better. But I really tried to stick through it. So I asked it to like sort of show me create more visualizations. I kept asking questions like, "Oh, like hey, why you know like like what do these different things do? Like what is a vector scope?" like yeah ask giving me like some sort of like it pulled literature elsewhere from you know from color grading and put it in here and then finally it made showed me a bunch of examples of like okay this is you know your old version this is a new version ultimately color grading is kind of like a shader it's like you know you take in a pixel you output a different pixel color and this is like a good mental model for me to build I knew what a shader was and the thing that really got me here was like I was able to build a visualization where I could go over every pixel and see how the value would change over time, you know? And so like I think through this like visualization, I was able to figure out, okay, what does color grading do?
19:01 And then I could ultimately tell Claude like, hey, I want something like this. And like the insight I got here is that you want to grade your skin differently than the background because humans have a larger lower dynamic range for a skin color. Like it'll look weird if you're like kind of purple, but the background can be kind of purple. You know what I mean? Yeah. Yeah. Yeah. And and in this particular case, my skin color and the background were very close together.
19:30 And so there was like some, you know, more interesting work for Claude to do, but like I was able to like express this problem and sort of figure out, okay, what what is wrong with it, you know, like what could be better? and then like fix it, right, through this like process of like grounding down my unknown unknowns. And yeah, I think that like there is so much work to do in that case, right? And I think like some of what we, you know, I think we sometimes oversimplify agentic engineering where we're like it's just prompts, it's just loops or whatever.
20:00 And I think it's more like no, we need to know what we're doing. We need to like learn more and then like be better at prompting and and make sure we're creating like you know valuable work, right? And in this case it this video and would not have happened if I hadn't been able to you know to to edit it myself. it would have like we actually couldn't pay it someone to edit it fast enough like we just couldn't. so yeah there are some techniques on like how to stay in the loop.
20:29 these are like you know I I'm writing more on this like I think exploring with Claude and brainstorming asking them to interview like we talked about in the last talk building technical plans explanation implementation notes explainers I'm not sure if I don't think I want to go one by one over these things. I think it's just more like the highle idea of like you probably have a lot of unknown unknowns, right? And you're probably not being ambitious enough kind of as a result almost and like how do you sort of free ourselves of this so that we can now like unhobble Claude to do more more useful work. so yeah that's that's that's my talk. Yeah.
21:14 >> A question here. Yeah. >> Claude tag. >> Yes. >> How much of your workflow has shifted over to cloud tag? >> quite a lot. but I think like I still do a lot of like exploration with cloud tag for example. I like sort of do a lot of thinking through it. So I might ask it to generate like an artifact or something and view it on my phone. but most of the like in the loop coding still happens with cloud code. and it's more like you know cloud tag is great at getting work started. It's great at managing jobs and things like that. Great at being multiplayer. Yeah.
21:50 >> Is cloud tag still single player mode within the organization or did it is it multiplayer now? >> Oh no, it's multiplayer by default for everyone >> by default but has it been adopted as multiplayer like organizationally? >> yeah I think so. Yeah. Yeah. We have it in feedback channels and things like that. people do together. Yeah. >> Cool. >> Yeah. >> could I get to know a little bit more about your process for like empathizing with the model and then designing the environment around it because it's not like pure in a human sense like >> if you're leveraging pure reasoning then you want like more composability or like >> you know some level of common denominator across every tool rather than like adding on more and more.
22:24 >> So is there any other like intuitions that you kind of learn through like >> iterating with different tools? Yeah, I I'm trying to figure out yeah, I I think there's like I I overall think this is an art kind of right now. Like I I think that one of the things Okay, I do need to Oh, actually actually no, I I skipped this part of the talk. Maybe I can actually go back to this. yeah, some some examples of how Claude gets better over time. I do need to take credit for this actually. I did the interview grill me thing like I you know I think Matt later made it into a skill but I was the first one to like sort of identify that. but ask user question is a tool that I built in cloud code. It helps cloud code ask you questions and th this how you I've used it has changed a lot over time where like you know initially the model could just call it once and then I tried being like oh can you chain it together? Can you interview me? and then you know would like put like 30 or 40 questions together. Now I ask you to build HTML reports you know and and sort of like get much more in depth and then select the answers from the HTML report.
23:32 and so the how you get through this progression like Opus 4 could barely call the ask user question tool. It was like actually kind of hard. and now it's like there's so much capability overhang on top of it. similarly with like markdown and HTML like this is something that you know I've talked about a bunch before where you know originally markdown was a way that like maybe you know the model kept itself on track and now then it became a way of communicating to you and now like it can create you know really rich HTML reports and this is like the like the model is getting smarter but in these spiky ways right you need to switch formats from markdown to HTML in order to unlock its capabilities. How do you figure that out? I I don't have like a science for you. You know what I mean?
24:17 Like I think that's like I think the maybe trillion dollar question. >> So just to clarify, there's like two different ways. There's one versus environment where like ideally it's just like bash or super lenient where it's like you have all these emerging properties that come out of just like more reasoning. Yeah. >> And then the other is like it could only work with the information it has. So that's like more signal and like having these or necessity these tools in the first place.
24:37 >> Yeah. Like the context and things like that. Yeah. I I mean I I that seems right but I need to think on it more kind of if that's the right paradigm. Yeah. >> Yeah. >> question. So to me a lot of agentic work seems a little bit like the old school unsexy waterfall like software development right where you have specifications then you have implementation QA all of that right? >> Yeah. So how would you say one could go beyond that specifically for agents because right now it's like interview me that's basically requirements and specifications right like iterate that's QA >> sure >> so how do you think one can make it better for agents >> what's the end goal you think like is there like a problem you're seeing >> build software right like like for example like I've been using it a lot for software building >> so a lot of times I have an idea of what I wanted to build and I iterate, you know, via different ways.
25:36 >> Sure. >> So, but I still have somewhat of a product in mind. >> Yeah. >> That I want to build. So, I kind of follow this process, but I'm trying to figure out how to do it in a better way. So, that maybe because like to your point, there's a capability overhead. Like, what am I missing? What am I not thinking? How should I think about interacting with agents in a more sophisticated way? >> Yeah. So the the problem is that the agent is not building exactly what you want and like there's some mismatch between you and the agent. Is that right?
26:03 >> It's building what I want. It's just it takes longer than I would want. >> yeah, I mean I think there is it kind of depends on your particular flow. Like I I think like one of the things I do a lot is I build a lot of prototypes first. And so these prototypes can be really cheap. like you know you can sort of like build a mockup in HTML so the agent doesn't need to fully implement it and so then you can sort of get a sense of what you want. I often find that I don't know what I want when I'm building something and there's like a iterative process of finding out what I want that can happen like quite deep in the implementation process and that's usually what slows me down is like I'm like 80% of the way there I'm like no that's that's wrong you know so like how do you like get that earlier and earlier and like what I like to do is I I have this like iterative like spec interview prototype see if I like the prototype maybe even like make a prototype PR are kind of all in along the way. then I might reset and take those learnings and and start again, you know. because yeah, I just like I want to build the most valuable thing I can, you know. Yeah.
27:13 >> Yeah. Daisy. >> Yes. >> Do you still use plan mode? >> No. Yeah. We need to do something about this. Yeah. Yeah. Yeah. Daisy also works on cloud code, so Yeah. Yeah. Yeah. >> yes. >> All right. Last question, Robbie. >> how do you decide what to like have Claude spend like take time to learn from Claude? I feel like my primary form of brain rot these days is like Claude explaining stuff to me and like an hour later I'm like, "Oh, I need remedial matrix algebra to understand this." Like just go home. I don't know.
27:44 >> Yeah. Yeah, I know what you mean. Like, let's see. This is a good question. Honestly, like I don't think there's like an easy answer. I I think that like you know, like Yeah. Like ultimately it's like what is the goal that you're doing? I I do think one thing about agentic engineering is that it's so fun. Like sometimes you can just be like just do this and it's just fun and you're not actually driving an outcome. I I'm not saying that's bad but I'm saying like you know like sometimes you do want to be like okay what what is wrong about my output? How do I get better? And and probably the answer is that there are unknown unknowns you have, you know, like there's something that you're like not able to express well enough, you know, or maybe you don't know enough about the user or or whatever, right?
28:26 And like how do you like answer that? And hopefully Claude can help you answer. but yeah, you know, it's yeah, you know, Claude, you can learn about everything. Not you don't need to learn about everything, but there are particular things that you need to. Yeah. >> Cool. Thank you. >> Thanks. >>
Summary
- The speaker highlights the need for deeper understanding and ambition in AI development, suggesting that many users underestimate the potential of AI tools.
- They introduce the concept of "capability overhang," where the actual capabilities of AI models exceed what users currently utilize.
- The importance of human-agent interaction is emphasized, suggesting that effective prompting and communication with AI models is a skill that can significantly enhance productivity.
- The speaker shares personal experiences with AI, illustrating how they learned to better utilize AI tools like Claude for tasks such as video editing and coding.
- They discuss the challenges of unknown unknowns in AI, stressing the need for users to identify gaps in their knowledge to improve their interactions with AI.
- The conversation touches on the iterative process of building and refining projects with AI, advocating for prototyping and continuous feedback to align AI outputs with user expectations.
- The speaker encourages a mindset shift to explore the full potential of AI, suggesting that users often limit themselves by not fully understanding the tools at their disposal.
Questions Answered
What is the importance of being ambitious in agentic engineering?
The speaker emphasizes the need to recognize unknown unknowns in agentic engineering and the importance of being ambitious in learning and prompting to create valuable work.
How do AI models improve over time?
AI models do not improve in a linear fashion; their growth is often unexpected and can emerge in surprising ways through experimentation and exploration.
How can we uncover the capabilities of AI models like Claude?
By actively engaging with the model and asking specific questions, users can discover unexpected capabilities, such as video editing, that they might not initially consider.
What strategies can be used to address unknown unknowns in AI?
To tackle unknown unknowns, one can ask detailed questions and seek visualizations to better understand complex topics, such as color grading, which can lead to deeper insights.
How have AI tools evolved in their functionality?
AI tools have progressed from basic question-asking capabilities to more complex tasks, such as generating HTML reports, showcasing their increasing sophistication and adaptability.