Transcript
0:00 I was excited to sit with Hamel Husain, the founder of Parlance Labs. He walks through why you need to be moving fast and experimenting with AI, but also why you really need to slow down and ask yourself the right questions. Talks through harnesses and guardrails you need to be putting on, and ultimately why it's really important that we're digging in, but also stepping back and asking ourselves, "What should we be doing here?"
0:33 Super excited today to be joined by Hamel Husain from Parlance Labs. He is in the thick of it. We talk about companies implementing AI, being challenged with AI, how they're going about solving it. He is both a leading a consulting firm, well-known speaker, and author on the topic, and I'm very excited to have him here to talk through his thoughts on this crazy world we're living in today. So, welcome. Happy to be here. Thank you. Before anything else, can you just introduce everybody to Parlance Labs and give a little bit of your background?
1:06 Yeah, a little background of myself. I've been a machine learning engineer for over 25 years. I worked at a lot of startups, and small and bigger ones. I worked at Airbnb, GitHub, a bunch of other machine learning startups. Been doing this consulting, independent consulting, like I'm an independent developer. I've been an independent developer for about 3 years, helping people build AI products. And the bottleneck that I kept seeing is people are struggling on how to measure and test their AI products beyond vibe checks. And so, I decided to focus on that. And so, I also do training and education on the subject. I write books.
1:44 I'm writing a book with my co-author Shreya Shankar on AI Evals. And I yeah, I talk a lot about Evals. So, just for the jet for general population, can you describe what an Eval is? It's kind of a hyper-loaded term, but basically what it is is how do you do data analysis and debugging on your AI application in a structured way so that number one, you know what's wrong, but then also you know how you should prioritize, what you should fix, and how to design metrics to measure things, especially when you have stochastic outputs. Like the outputs of AI are stochastic. They're like text.
2:24 You don't really have There's not like a clear right or wrong answer you can deterministically test. So like how do you go about testing an application like that? That's what evals is all about. So obviously that's super critical. When you're beginning the process of implementing, do you need to define that or is it completely iterative as your needs in the model need models evolve? Yeah, so the way you start with the evals is data analysis. And what you do is you do you go through a process called error analysis, which is a kind of data analysis that's kind of blend some qualitative analysis with some quantitative analysis.
3:05 And you figure out what is broken in your application. And based on that, you can decide what is the right things to measure based on what's actually happening in your application. So the thing that's different about software testing and AI testing is when it comes to AI products, there's an infinite surface area of what can go wrong. And so you need to kind of prioritize like what to measure and how to measure it. And so that upfront data analysis is key to help you zone in on, okay, like what you should do. Um so it's a little bit like this bottoms-up approach is really important to complement like a top-down approach of like, hey, um I want to I have some things I do want to test or I'm worried about certain types of failures. Um but I think people overly focused on the top-down uh to their detriment and they get lost in generic metrics. Can you dive into that a little bit? What What do you consider a generic metric in this case? So, if you Google evals or you Google how do I evaluate my AI application, there's a high likelihood you will stumble upon some vendors or some tools that will promise to or completely automate the testing of your AI application.
4:21 And what they'll promise you is they'll throw up a dashboard that has a bunch of scores, like helpfulness score, conciseness score, toxicity score, coherence score, you name it. And you'll get a beautiful dashboard with a bunch of metrics on it, usually on a scale of 1 to 5 or 1 to 10 or something like that. And what ends up happening is no one really knows what that means. Also, those kind of generic things usually don't correlate with what's important for you to focus on or to fix. So, actually those things are actively very harmful because they distract you and they make you burn engineering cycles just looking at metrics that don't matter. And so, what you need to do is you need to be very thoughtful about what you're measuring and make sure that the metrics that you do have a matter.
5:13 And you kind of have to put your data science hat on. Um it's kind of just very similar to product analytics. Like, you wouldn't take your product and throw up a dashboard with a bunch of generic metrics. You would think carefully about the way your metrics are calculated and if they make sense for your business. So, the kind of the same idea applies here. I have so many follow-up questions. So, you're beginning an an implementation process.
5:37 In your mind as an executive or as a leader, you know the outcome you're trying to get at. In most cases that of people I've spoken to, they sort of assume that the models have X% hallucination and and are wrong and are trying to close that gap by training and training and training and then having some human oversight. Is that a fundamentally flawed way to go about the implementation because they're not even setting up the right testing plan up front? So, if you go into trying to test your application with this idea that you need to reduce hallucination and help and you need to increase helpfulness and you you know, this is kind of like a generic metrics mindset.
6:21 >> Yeah, totally. And what that means is you don't really know what's wrong. You're just kind of going through some motions and you're going to end up wasting a lot of time and you're going to get lost. And it's a very appealing thought that hey, you can just don't worry about this testing stuff. Just plug in this framework and we'll calculate a score for you and you'll be fine. That's absolutely not the case and that's why a lot of people struggle. Um you know, frankly, that's why my business exists. If people weren't getting confused and they weren't getting led astray, then you wouldn't need my help. Unfortunately, I think people do are trying to look for the easy button. Unfortunately, in this case, an easy button doesn't exist. It's not the case that you have to do everything manually or you have to it has to be painful as you know, you can use coding agents to help you implement the evals and write some of the evals and wire up the plumbing.
7:15 But you have to be thoughtful about what you're measuring. And really that one of the big parts about evals is it's not just purely testing. It's a process that where you look at your data and you understand what good looks like. So, most people don't know what good looks like including myself, really, including anybody until you look at the outputs. And you what you need to do is iterate and kind of specify like, oh, like this is good and this is bad. It's kind of impossible to do that without in It's like an iterative way of looking at things. That's, you know, this like And that's part of the annotations that you might do with evals is like, you know, and that's what happens when you're trying to transfer your knowledge to the AI.
8:02 So, it's really hard to build a good AI product without doing that exercise. You know, maybe I maybe cuz I just came back from Human X where everyone's talking about how agents will do everything. Is there a world in which the automation itself can do the eval and then automatically fix it or is it inherently, you know, a human-to-model training issue? I think AI is really good at fixing bugs, deterministic errors in your product. Like the code's not working or it's doing the wrong thing or it's it's failing a test. But, AI cannot read your mind.
8:37 It doesn't know what you feel like good is. So, if a customer's interacting with your product and it's, you know, the product is not doing the right thing, it may look like it's helpful on the surface to an AI, but if you put your product hat on, you're like, you know what? We could have done better here. The inter You know, we should be able to help the customer in a better way or the help the user in a in a better way than this or, you know what? This interaction doesn't make any here.
9:10 Um we need to, you know, fix the tools we have. We need to fix our retrieval because we're not really bringing in the right sources here. And so, the AI doesn't know what it doesn't know. Like it doesn't have the ability to read your mind and sort of elicit what good looks like. What AI probably can do is walk you through the process a bit of evals and say, okay, like and kind of interrogate you um a whole bunch." I think that's where the future is is to say like, "Okay, let me guide you through the end-to-end process of evals and let's write it together."
9:48 Which is a little you know, that's kind of what good eval frameworks are trying to do. But you still need to have a human in the loop. When I hear what how you describe it, it you know, Claude Anthropic you know, has this skill builder for small businesses and individuals. They launched where it's very back and forth. Is this what you want it to look like? Is this what you want it to feel like? Maybe. I don't think it's going to be purely chat.
10:11 I think it's needs to be a bit different than that because what it involves is looking at lots of data. And you want to render your data in a very domain-specific way that's specific to your products. So, for example, if you have images, you need to render those images. You have emails, you need to make it look like an email. If it's a chat, it needs to look like a chat. If there's metadata that's involved with making decisions, you need to render that metadata in a very nice, easy-to-see way alongside other data so you can make a quick judgment on what is happening with the AI. Yeah, there's probably some software.
10:45 Um it's a little bit beyond just chatbot. But you know, I don't think there's anything really special. I think like if you zoom out a bit, I don't think agents are going to completely build software either. They will build this They will build it to your specification. But what is your specification? Like, you know, um the more non-trivial your software is, the more you're going to have to inject your taste and your specific point of view.
11:14 And if you don't have a specific taste and point of view, then your product is probably You know, it's not going to be differentiated. Or like, what are you even building? And so, I think that in the same way, uh evals is the same way. It's really an extension of that. Cuz like, what really what you're getting down to with e-vals is it's a process of like eliciting your specification and then measuring against that. So, in this world now where, you know, we like to joke that boards call a CEO, they're like, "What are we doing for AI?" The CEO calls VP It just you know, layers down to let's get something out and let's show something.
11:51 What goes wrong as part of this process when you think about how, you know, when you step in? So, the first thing that goes wrong is not even e-vals. The first question I ask is, "Are you using AI?" And are you using AI in a deep way? Are your engineers coding with AI? Are you building things with AI that's beyond just like copy and pasting to ChatGPT? If the answer is no, then I then you're not going to be successful building AI. Because you won't have a good intuition on what is possible and what's not. You're going to have really bad specifications.
12:27 You know, you won't kind of zone in on what is a good idea and what's not. You won't have a good mental model. So, I think that's the first failure point that people face. And the second failure mode that I see a lot is reaching for complexity too fast. So, okay, you want to build an AI product. Don't off the bat go for the most complex architecture and setup. Like, don't go for Don't like on day one reach for the most complicated orchestration framework and a graph database and a multi-agent thing. Start simple and build your way Build your way up incrementally. I see a lot of teams that just go straight to the most complicated architecture, um and then, you know, they can't really reason about what is happening and they don't have a good mental model.
13:16 Um but I don't I I think that last point is not necessarily unique to AI. That's always been a problem in software to some degree. In AI, it's a little bit more acute because you can swallow complexity a lot faster because of AI agents. So, how do you balance the sort of need to go slow with the reality of FOMO and pressure from all the different parts of the org. Like, if you're you know, director or a VP, what do you What guidance do you have for how they can manage up in these cases? It's not necessarily going slow. I wouldn't go too I wouldn't go slow. I would just be very deliberate. I think it's really important to be building with AI, be trying lots of ideas. Now, there's a double-edged sword to AI is AI allows you to build the wrong thing faster.
14:08 And build many wrong things faster and with more scale. But, it also lets you build the right thing faster. It also allows you to experiment with a lot of ideas. It, you know, it sounds good on paper like, "Okay, you can experiment with a lot of ideas." But, if most of your ideas are bad, then you're just going to sink in the bad ideas still. You're going to churn through lots and lots of bad ideas. So, you need to be a little bit deliberate. It's good to have good product sense and some good grounding and taste um in what you should build so that at least, you know, you're moving in some productive direction. Because, again, like agents can be very distracting on their own. They can be fun just to build. I think it's not really moving fast, just being deliberate in my mind. And then, trying to like do experiments and sort of double down on what's working and be very methodical about that. That's super interesting about AI being able to build bad things quickly.
15:11 In the last 2 years, people have gone from I'm very actively managing my AI, you know, to I'm going to let them write this. I'm sure you've seen You get these emails where it clearly nobody read it over. They're just Somebody put it in AI and sent it. When you go on an organizational level, how big a risk is that, do you think, where there's just this lack of ability or lack of desire because our brains are getting trained to not push back and not actually ask the questions that would lead to real deliberate implementation? I think it amplifies who you are.
15:43 So, if you are a mediocre person who is okay with writing slop, for example, um it's just going to amplify you and you're not you're going to be more susceptible to say, "You know what? You know, I don't mind." Even if you're not a mediocre person, you can still everyone gets lazy and gets tired and can get trapped in this, but I think it does amplify your natural tendencies. You know, just statistically, there's more average people than above-average people, and I think it you will see it like this amplification of slop um if everyone adopts it because it takes away all the friction of doing the work in the first place. You Before there was some friction between you and you know, pushing this information out.
16:26 Now there's no friction. So, that's an unfortunate kind of reality of it. And it's really the same thing about it's not even just like yeah, it's not just about building the wrong things. It's saying the wrong things. It's uh doing the wrong things. It's doing it you know, it just amplifies you it could amplify you in the wrong direction like very fast. No, I'm kind of hopeful like maybe we will find ourselves. We'll find a way to deal with it. We'll figure it out. And I do think that people do notice the difference. Like people can smell AI. You know, they they understand when And you know, I I I think people they'll we'll adapt in some ways. Maybe not completely, but there'll be some adaptation. We won't be completely lost.
17:07 So, you've just described a chain of events where things can go really well or mediocre or really right. When you walk in and you start an engagement, like what's the most common problem that you run into? One is swallowing too much complexity, which I already discussed. Another one is there's a mandate from on top that hey, we need to use AI. And people take a existing product surface area and just slap a chatbot on it and say hey, we got AI. You know, that is a very mediocre experience that kind of doesn't move the needle. And so, I think you to yeah, to build AI, you have to you should be pretty you should try to be thoughtful of how you can actually help the user. Um like accelerate their workflow and do things faster. For example, instead of putting a chatbot on your product, consider putting an MCP on your product.
18:04 Or consider exposing APIs for your product. That is likely to be way more helpful to a lot more people than just putting in a chatbot on it. People are reluctant to expose APIs and MCPs on your product cuz they feel like they're getting disintermediated by the AI. Now, you're not even opening the application potentially. You're just chatting with Claude Code or ChatGPT or whatever to interact with your product. But, at the same time, you know, I don't think people are going to be clicking on menus and clicking buttons for that much longer. I think it requires a little bit more and this comes back to are you using AI?
18:37 Cuz if you're not using AI uh deeply, then that's where these problems stem from cuz you don't have a good mental model of where the puck is moving. That's super interesting. Can you just describe like a a use case or a case study chatbot versus MCP because we just internally at my company, we switched providers because they didn't have an MCP for Claude. We just canceled our our whole account. But, I I think that nuance is really lost and and I'd love to hear a case study or examples that you might have.
19:11 >> like a really big example that I think everyone can relate to is Google. So, recently, up until a very recently, it was difficult to interact programmatically with AI with Google works like Google Workspace, like Google Docs, Google Sheets. They had integrated chat with Gemini, but it was very poor like to do anything. You know, to like modify a spreadsheet, to modify your calendar, to look through your email. It wasn't really great. It was very frustrating, honestly. Um and then eventually like they exposed the Google team exposed like a Google Workspace CLI. Um there's some other third-party stuff um to make these more agent-friendly, and that makes a huge difference. That's what everyone uses now. You know, like if you're using an open claw, for example, you know, it's using if you want to wire it up, it's using that. So, you know, that's that's really huge. Another example is that may hit home is if you just slap a chatbot on your product, you have to think really carefully about the interface. For example, if your chatbot is scheduling a meeting, you don't want to go back and forth on just text. Like, "Hey, um here are the meeting times." and bullet points.
20:23 Meeting one meeting time one, time two, time three. And then you have to go, "Oh, yeah, the 4:00 works." That can be very brittle. You should expose an interface. Like, "Here's a widget. Select one of these times." Great. Let's select the time. Great. Uh confirm. That makes a lot more sense. But people need to think a bit more holistically about like, "Hey, like what is the right interface? What's the right workflow here?" where the user can get visual feedback and be confident that it works. And that you can also avoid bugs. Cuz if you're just trying to do everything with text, like just pure chat, then I don't know if the tool fired correctly. Um you you know, you don't know if the tool fired correctly, either. Maybe it did, maybe it didn't. Um I want visual con- confirmation. And you, as a person serving that, you want it to work more deterministically. So, it's so not everything needs to go through an LLM, right? So, it's like how do you have the right approach in product thinking? Is part of the eval process putting in guardrails to make sure we stay within certain harnesses or frameworks?
21:26 Or where does that come into play? Okay, so the question is guardrails. Okay, so first let's talk about what a guardrail is. So, guardrails are very specific kind of eval that sits in between the request response path and blocks a certain output from being shown to a user. For example, you don't want your AI to talk about competitors or maybe use profanity or something like that. So, you just want to block it. My blocking it means either you just prevent the AI from saying anything or you just make the AI say something generic like, "I can't help with that."
22:00 And we've all seen that. So, a lot of people think of guardrails, they also think of off-the-shelf stuff. They're like, "Oh, I'm going to go to this framework and I'm going to get the profanity guardrail or I'm going to get the um unhelpfulness guardrail." And you have to be really careful cuz if you use a guardrail off-the-shelf, you're just using someone else's prompt and it's very likely that that someone else's prompt is not going to work well for you.
22:26 I've I've actually uh looked at these prompts quite a bit. Those prompts have examples that are very specific to certain domains. Often times it's like a shopping domain or a travel domain or something like that. But if you have a if you're doing something illegal, you don't want a travel domain example in your guardrail prompt. It's not going to it's it's not the it's not the best fit. And so, you have to go through the like a similar process of evals and decide, "Okay, like which failures you want to prevent against?" And you want to prioritize failures that actually happening or failures you can simulate. If you can't reasonably simulate the failure or you don't see it actually happening, then it's kind of a lower priority. But yeah, guardrails is a special type of eval. Usually you want it to be fast because it's blocking the response or it's in the path. But it's really very similar to other kinds of evals. It just has that these additional characteristics. So one layer of complexity on top of everything else is most companies are trying to figure out the stack that they're using which then manages for cost, for model usage, and as you're trying to create a testing and eval plan, but also trying to manage for a changing world where you're switching models.
23:42 How do How do you recommend people think through that process? Yeah, so the best thing to do is to use the most powerful model you can to start with to make your life easy. Just pick one to start? Yeah, pick one to start. Um if you can use like, you know, the best open AI model or the best Anthropic model or something like that. Something easy. Hopefully it's something you're familiar with. So again, something that you're using to build stuff already.
24:09 You're already using in your coding agents. So use that cuz you might already have That is important to like, benefit from that intuition you already have. And then what you should do is build an eval harness with the metrics that matter. Then you can try other models. You can try backing off to smaller models, different models, and you can sort of see, okay, what the tradeoffs are between latency and cost and performance. And sort of reason about it from there.
24:36 The The reason why it's useful to start with the most powerful model and then back off is it lets you to build more simply to begin with. You don't have to Yeah, you can try to see like what is works in the search for the most easiest case of trying to get the AI to do something. And then you can back it off and see what happens. And then you can more reasonably like reason about the trade-offs. Whereas, if you start try to start in the other direction, sometimes like you have to build some more complexity to get to the same performance. And you don't know it's not really Yeah, like and then you're you kind of stuck with the complexity. But you're only just as I process everything that you're saying, in this world most executives that I talk to, like oh, we're going to go in and we're going to cut all these costs, and it's going to be fantastic. But what I'm hearing you say is you start with a process, there's all you need a layer of people. In order to really maximize it, you you're going to continually need watching, evaluation, building, experimentation, which would sort of speak against any real cost efficiencies in some parts of the org.
25:41 And you wrote about this, you just wrote about data science, actually. It was this you know, the view that data science is going away, but actually what you're describing is a in a world where data science becomes more important than ever. I do think there's going to be a lot of cost reduction, for sure. I don't think that roles are going to be wholesale eliminated completely. Like there'll still be a software engineer, there'll still be like a product manager, still be a data scientist.
26:06 Maybe those roles will collapse into fewer people, so people will wear more hats um than they than they have before. They're able to span more surface area. So, I do think that costs will decrease. You know, like you don't need as many data scientists as you did pre-AI. But it's still good to have the skill somewhere in your organization, even with AI. Cuz like yeah, what a data scientist is doing is they're asking questions. And the ability, as you know, the ability to ask the right questions is directly proportional to the quality of output you get with your AI. I do see that there'll be a drastic cost reduction. Now, I guess it's left to be seen like what direction you take as a company with AI. Like do you just try to hold the line and do more with do the same with less?
26:59 Or you try to do more a lot more with the same people. I'm currently I mean I'm kind of in with the view of like, hey, you have to do more otherwise you're going to die in a lot of cases. They might They might be certain cases where yeah, you could just stay the same and just do it like it makes sense to just be more cost-efficient. But for a lot of tech related things, um yeah, I think growth is important. A lot of the model companies, the OpenAIs, the Claude I think just announced last week, they're trying to set up their own implementation teams to help move directly into I'm going to work with private equity to go ahead and change how you think about AI. What's your advice for going directly with a model company versus using an implementation firm versus using some hybrid? And how should people think about the process of who's helping them think through these changes? I think the question is about when should you rely on a third party or get the help of a third party when building your AI applications? And fundamentally it's it's very similar to should you engage with a consulting company to build your AI? And that in turn is kind of leads to the question of is AI competency in building this AI product within your core competency? Like is this something that's important to your business? You know, so like if it's a software business, it's hard to see how building an AI product is not within your core competency. You know, if you're exposing a product of any kind that is software, yeah, it's really it's really hard to see how it wouldn't be. And I haven't yet encountered a company where it isn't in that where I feel like it isn't in their core competency. The reason is is because if you're doing knowledge work, fundamentally AI should be within your core competency. So, it's very difficult for me to think of a situation where AI So, you should be really careful about getting a consulting company to help you implement AI. Now, where I think you should maybe you can maybe get help there is to upskill. You should absolutely upskill your team and so they're not dependent on a third party. Being dependent on a third party, especially for something as important as AI, is very risky in my mind and it's also and uh kind of a an anti-pattern.
29:17 Because what are you going to do when those external parties leave? You need to make sure that you are like you treat it as a training exercise, like a deliberate training exercise, and you're not using it as a crutch. And when And when you go in and talk to companies, do you have a framework or sort of a a roadmap that you show them that takes them from point A to point B and then allows them to wash, rinse, repeat on their own, basically? I put companies through sort of a boot camp where I put them through this Evals course. First, I make sure that they're in the right place to do Evals, meaning they're using AI internally, they already have an AI product, and they just they're at a place where they want to make the AI product work really well.
30:01 And they want to know how do we test it? How do we measure it in a way that makes sense? So, once we get past that, what I do is I have a course that teaches people Evals. So, I give my clients access to that course. And at the same time, what I do is we go through their data and we debug their product and we do this whole end-to-end Evals process, but we do it together.
30:24 You know, we basically pair program and build the whole Evals end-to-end. And I they're basically doing it. I'm telling them how to do it and how to get them unstuck according to their data. But by the end of it, they don't need me because I transferred all my knowledge, given them all the tools they need, and then we've gone through some prac- we've done a bunch of practice. Uh so, it's kind of like going to driving school, I would say. Yeah. Cuz like, "Hey, like you might do a little bit of study of the rules, but then I'm going to get in the car, and we're going to I'm just going to drive with you until you can drive."
30:56 And then when we're done, you can just drive on your own on your own. What's interesting is this argument I keep getting into, which is I'll have this discussion with someone, and then they'll go to to X or Twitter and read a quote from somebody at OpenAI that said they wrote 10 million lines of code all by themselves, and they'll skip your example and say, "I don't need any of that. I'm just waiting for self-driving car." And and how do you can convince people that that might not be the right approach? We can use like analogies all day. Yeah. So, just to humor the analogy.
31:26 >> Yeah, yeah, yeah. You have a self-driving car, but you got to tell it where you want to go. Right? So, so, um that's really what this is. Like, you know, fundamentally, you know, you it's really, really difficult to tell to know like what to tell an AI to do and keep it on target, keep it on task without a harness, without uh an environment where it can test itself, and it can get feedback on whether or not it's doing the right thing, and keep itself in check.
31:57 Um and so, with these million-dollar or sorry, these uh million lines of code things, blog posts, they're all backed by this like harness engineering idea. Uh that's the only way you can do that. And inside this harness engineering are metrics and logs and traces and observability, and that's keeping the AI on track. So, if you want to have a really good harness, you have to have evals. So, evals is a huge part of the harness. It's almost all of the harness.
32:26 It's interesting cuz in the world of social media, those nuances get lost. Yeah, definitely. Yeah. You know, even I feel like I'm informed enough to to and I read that, and I'm like, "Oh man, it's the easiest thing in the world." But it it clearly isn't. I'm implementing AI today. I'm deep in. I'm experimenting. I'm kind of living what you described. Should I just expect that I'm going to make a ton of mistakes over time and some portion is going to be money's going to be lit on fire um and that's part of a adapting my org to this new world?
32:58 Yeah, I think so. I mean, I think it's hard to do anything without making some mistakes. Uh I make mistakes all the time. And I think it's uh you have to be willing to tolerate some mistakes if you're going to experiment. And you're going to especially with AI that's moving so fast. I think the main idea is to have a very experimental mindset and to encourage people to be using AI at the frontier as much as possible so they can have those mental models so that you can make less mistakes cuz you know the biggest mistake is building the wrong thing. You can eval you can eval the wrong thing all all the way to hell.
33:37 But it's not going to help if you're still building the wrong thing. Like it doesn't matter. Yeah, it's just really useful to to experiment a lot. Well, it's interesting when I hear you say that it's most companies that aren't in the tech world don't have a quote-unquote R&D budget. But it's almost like maybe every company needs an R&D budget now. Yeah, I mean, and so what I mean by experimentation is not this idea of this expensive lab with these like supercomputers or beakers and flasks and all this you know like chemicals and like it's like, "Oh, this like fancy R&D." No, I'm talking about you sitting at your laptop using cloud code trying to build some stuff. You know? And so I think you don't need an R&D budget. Maybe need a token budget.
34:20 Uh but not really. I mean, you can have a $100 plan and get pretty far and you can get a lot of intuition by using these things very deliberately to to solve problems. Uh just shifting gears for a What's your personal tech stack out of curiosity? Personal tech stack? Um so yeah, I use Claude Code a lot. I use CodeX as well at the same time. I'm constantly experimenting with stuff. I use Cursor a little bit. I don't have I used to have opinionated tech stack. I used to you know, I used to be a Python developer just like only write Python.
34:53 Uh mostly data science stuff, but now I'm like all over the place. I don't really care. I use the tools that the AI wants to use. So I do that. I have a bunch of skills. I have a bunch of tools. I you know, try to slowly build my tech stack. Um I experiment with stuff constantly. I'm experimenting with Open Claude kind of counter to let's say the narrative. I haven't found it to be super useful yet. In my company I have it's me and I have a bunch of other kind of talented developers that are either my friends that are in my Slack channel uh that work at other companies.
35:30 Um and honestly, we spend most of our time with Open Claude we've been spending most of our time improving Open Claude or like fixing Open Claude or building tools for Open Claude or building tools that can build tools for Open Claude. And then we step back we're like, wait a second. This is all fun and amusing, but like we're not actually doing anything. And so um it is a realization that I came to. And it's like you can do like all the same kind of you know, recurring schedule scheduling tasks ambient now through like Claude. You know, they have like dispatch and Claude co-work and schedule tasks and stuff. You can do a lot of stuff there. So I don't know. It's it's but you know, despite that we still experiment. We're like, okay, like is there is there like a place where is interesting or useful you know, like always try to find out just by using it constantly. So yeah, I was just using stuff.
36:25 I had the quintessential Open clock experience where I set it to manage a marketing campaign and it decided that it was working so well it up the budget by like a 100% 100 times and I went from spending five a day to 500. Okay, well was it working well? At five it was but you know the harness going back to your eval would have been okay once it's at five take it to seven then take it to 10.
36:50 And I just was like you go you go wild with your you know evaluation and it jumped it up to 500 and I figured out after a day but it was a good it was a good lesson in Okay, what is it doing now? Is it is it still running? It is but now it's like you know I've gone the other way cuz I don't pay that much attention to it so now it like goes up a dollar a day which is almost boring. So I need to there's there's a middle ground that I somebody could help me experiment with or I could spend more time on but had I set up the project intelligently beforehand this all goes back to what you described and done my own evals versus just jumping in and you know guiding I would have been much more successful. Yeah, what I've noticed is okay so even though AI allows you to do things faster and maybe span greater surface area to do one thing like super well you still have to focus a lot on it.
37:46 So in that way nothing has changed. So for example if I want an AI to be really good at video editing I probably need to really focus on video editing for a couple of months. Maybe exclusively and make it like really really good. As good as possible. If I just do use like someone else's prompt or someone else's skill or try to like vibe it out in a day I'm going to get some like very mediocre and it's going to still be exciting because it's going to be way more than what I would do normally but it's not going to be like, oh, this is amazing. It's not going to be a par with maybe like a really talented person, per se. But, you know, it's still useful because it's like free. But, so it's very interesting. It's like, yeah. It's funny cuz I just wrote this piece. You know why AI isn't great for the ADHD population? Because and and you alluded to this earlier, it magnifies whatever, you know, greatness or weakness you have. You know, as I talked to my most ADD friends, they're building 25 things.
38:51 And if I talk to my most focused friends, they're nitpicking the product to the point of who cares. And somewhere in the middle is the right answer. Um but, you have to manage your own personality in the world of AI. Well, look, thank you so much. Uh you know, Parlance Lab sounds like it's doing important work in terms of helping us all make it useful, which is the whole, you know, point that I'm interested in is how do we actually get this from it's a company mandate to it's actually helping my company grow.
39:20 And thank you so much for spending time with us. Yeah, thank you.
Summary
- Rapid experimentation with AI is essential, but it must be balanced with thoughtful evaluation.
- "Evals" are crucial for structured data analysis and debugging of AI applications, focusing on identifying and prioritizing issues.
- Organizations often struggle with generic metrics that do not correlate with their specific needs, leading to wasted resources.
- A bottoms-up approach to testing, combined with a clear understanding of what "good" looks like, is vital for successful AI product development.
- Companies should start with powerful AI models for initial implementations and then explore alternatives based on performance and cost.
- The integration of guardrails is necessary to prevent undesirable AI outputs, but they must be tailored to specific use cases.
- Organizations should focus on building internal AI competency rather than relying solely on external consulting firms.
- A culture of experimentation is crucial, as it allows teams to learn from mistakes and adapt to the fast-evolving AI landscape.