transcribe

Not Another AI Podcast Ep. 04: What happens when a company replaces dashboards with AI?

Ask Enola · 44m · transcribed May 2026
More from Ask Enola Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Transcript

0:00 where 80% of the company was going to dashboards to get their information to now 80% of the company goes to their favorite AI tool to get the information that they need. So it has Yeah, cool. Yet another episode. Let's start Paris maybe with your story of uh open door and how you started and love the arc.

0:39 >> Yeah. Uh well, first of all, thanks for kind of having me on the pod um and kind all the work that you do for the community. The u so a little bit of background. I'm Paris. I'm head of data at Open Door where I lead data science and data engineering teams. I've been here for about four years at this point now. Um and when I joined the company um the company ran on a decentralized pod structure which meant various business units had various analytics folks and data folks reporting it to them. Uh we had about 80 odd people doing data stuff at open door. Uh fragmented data tech stack. Every kind of team had their own tech stack. We had five different analytics tech stacks at the business. Three four different data warehouses and what have you. And so the common challenge that I saw when I was working with our executive team was there was lack of trust. Like the numbers that we're looking at, there was like five different versions of the number. Uh >> yeah, the finance tells me this, marketing is telling me this. I don't know what to trust. Let's not trust anything. Let's make decision based on gut.

1:43 >> Exactly. And that's something >> the talent bar uh also needed an upgrade because the decentralized pods that basically hired bunch of u uh uh less experienced folks and they were expecting them to be thought partners to executives which by design kind of doesn't work. So the data teams ended up essentially becoming a report function. Um and so what I >> was that part of the hiring uh newbies was that from a uh do you know if it was a cost cutting perspective like a keeping cost down perspective or what was why was that done?

2:19 >> I think uh mostly talent bar is not uniform. So uh what that meant is that some teams would hire for different kind of talent at different levels and they'd be calibrated differently and because of that you kind of end up in a scenario where if the person has not done data science or data engineering they won't probably in their first shot get that hire right and a lot of this hiring managers were hiring their first data people in their career. So which is why you kind of end up in that scenario in a decentralized kind of fragmented pods.

2:55 >> So they were inexperienced managers as the company was scaling and growing pretty fast. >> Yeah. Inexperience for data experience for their craft of course. So they're excellent at what I did just inexperienced for hiring their first kind of dedicated data person. So over the past four years I've kind of transformed the company to not just trust the metrics uh also consolidated all the various data tech stacks brought down our operational expense in terms of uh the cost that we pay for tooling but also significantly cut down the number of people that are dedicated to this craft. Uh so now right now we have a 25 person centralized pod um across data science and data engineering and we partner we have a hub and spoke model so we partner very closely with all the various seuite on the key decisions that needs to happen and over the past six months in particular uh we have changed the company from where 80% of the company was going to dashboards to get their information to now 80% of the company goes to their favorite AI tool to get the information that they need. So it has been a AI transformation journey that we've been on on the back of all of the work that we did previously which is the semantic layers, the data catalogs, the metrics definition. So all of the work that we did was leveraged to plan the last 6 months for the AI native transformation that was uh that happened.

4:20 >> Yeah, that's that's great. Shanti and I we've talked about foundation laying the foundation for AI in one of the episodes. Um and so yes you if you don't do that >> anything which comes after garbage in garbage out if you don't have a proper semantic layer or catalog >> whatever you build on top of it will not be useful. So I think that that makes sense and and also interesting that you went from 40 to 25. That's what I heard if I heard it right.

4:45 >> 80 to 40 to 25. Oh, 80 to 25 in terms of data data people, people who are in this data roles. That's a that's a huge one. Um, would you say is that mostly because of uh consolidation or is it more because of using AI and being more efficient like what's what's the story there? >> Uh, the 80 to25 happened before AI. So that was a function of consolidation of our data stack. But instead of five different analytics stack, you need one.

5:21 So that's efficiency boost. And also instead of like having four junior people pulling reports, if you have one data engineer who does the data model and one senior data science doing more analytics and experimentation and higher value work, you can replace kind of those folks uh in a more kind of thoughtful way. So that doesn't necessarily mean that the opex reduced in a similar way because we invested heavily in the high talent. So talent density over like talent in quantity.

5:51 >> Yeah. >> Was the key principle that we led with. >> Yeah, that makes sense. That makes actually to me it makes sense and and to also the listeners it makes sense for us to think about how the future is where the future is going. Uh at some point we had seen the wave where you know hire as many people and get this and like for this example customer support right like you need people or there was a thought process you needed bodies not now but you needed bodies so the dynamics there would be different versus here in in data team it's it's quite different um I'm also interested in this 80% going to dashboard to 80% going to AI tools of their choice so it's not a consolidated tool like what's what's happening Yeah.

6:37 So we have plethora of AI tools that we use but the majority of ones are folks going to the cloud ecosystem. So anthropics ecosystem of cloud code, cloud desktop, cloud co-work, that's what people use to connect to our MCP which is hosted on snowflake which is our data warehouse. Um but then depending on the team that you're part of uh there are other tools that we use as well. So we have a voice AI system that uses snowflake data as well. We have operational workflows on a tool called gum loop which connects to snowflake to get bunch of work done. Um and then we have an internal chatbot type installation uh which also connects to our data warehouse MCP and gives folks access to querying the data through that portal.

7:30 So there is a intentional choice to go and try different AI tools because the landscape is also moving way too fast that we are not thinking about consolidation just yet. We are thinking about empowering the users with their favorite AI tool and getting out of the way but putting harnesses in place so that when the MCP is called the semantic layer gives the same answer across all the tools because it's the same MCP that is being called by all these tools. So that's what we invest in and then let the front end be fragmented for now. I'm sure that'll get consolidated in the future, but right now because of the pace of AI tooling, we are leaning in to try as many things as possible.

8:17 >> That's something we've also seen where we'll build the backend part of the MCP, so the actual like how the tool works to make that somewhat deterministic. You ask for this thing, it's going to translate that into an actual query that's going to run. We've verified and vetted some of those queries so that we can have some data lineage and go, "All right, this is what the source data looked like. Here's the query that was run. Here's the calculation that was completed." But other front ends can reference that MCP and make it work that way. Build some confidence in the users that that's what they're they're getting good responses back.

8:54 >> I love the word deterministic that you use. That's essentially the name of the game, right? Like how do you make deterministic as possible? Uh so I love the >> which is funny because we take this stochastic process and we're like how can we make it less. >> Yeah. >> First we wanted it first. The only reason it this evolved is because we you know somebody decided to randomize >> the output or the weights and then now we have to come back to the mean like give me the right answer whatever that right is.

9:25 >> Yeah. and figuring out where you want that and how much play you want to give a model. So be like, well, I want you to choose which tool to use, but as soon as you pick that tool, I want the tool to work the same way each time. >> Yeah. And there's even difference in difference. I mean, I'm seeing differences between 4.6 and 4. Open opus 4.6 and 4.7 too. Yes. >> So even if even within the same model family, there are uh significant differences.

9:53 >> Oh, 4.6 6 in March is different than 4.6 in April. >> Oh yeah, that's story too. That's true. >> Um I'm curious about Paris when you said the MC the backend MCP models are the same. Are you guys building different MCP models built for purpose like finance or like if you need to think about customer support, this is what you go to or is it a singular single source of truth that everybody's connecting to? So there's a source of truth snowflake mcp but >> that calls our semantic models but every domain has different semantic model. So the MCP takes the prompt classifies it calls kind of the right query and kind of the right source and gets the user the data depending on the prompt. So the work that we did there was to train our semantic models to say hey if this types of queries come in these are the tables you want to use if these are the queries that come in these are the tables you want to use and we didn't stop at just giving you the tables we have converted our entire stack into code so our data transformation layer which is DBT is available to our semantic models our dashboard which also has thousands and thousands live SQL code in some cases is accessed by the AI mostly for verification loops.

11:27 And so if you think of your entire data text stack as a code and then give that code not just the text to the semantic model that has been the game changer because if we just gave it tables and some semantics that wasn't enough. it needed to be able to see the code behind the scenes for it to be more deterministic to Shanti's kind of language. >> Yeah. And I'm also I was also interested in your uh the little conversation we had about the verification loop because then when you first rolled it out then people started using it but they don't know if it's true or not. Uh so I'll let you let you share how that how that journey went.

12:11 >> Yeah. So initially when we launched it, we were all very excited. We had kind of within the data team ran like your usual eval right ask few common questions see what kind of AI comes back with we did that across our entire team and we kind of went through those I'm like okay this is super exciting it's kind of giving us the right answers data teams are also biased right because they know some of the numbers they can pretty easily tell whether that number is correct or not and because that's the world that they live in so we kind of launched it to our customers and week one was people getting used to it.

12:44 But week two in our kind of Slack channel, there was a spike in people basically saying, "Hey, I did this analysis through this AI tool. Can you just verify what the outcome is correct? It's urgent." And now the problem from like providing the infraction platform to how do we unblock these sets of customers that are now it's a good problem to have because now they're using the tool. There's nothing better for a data team to see that their work is being kind of utilized but it creates scaling issues which how do we now empower all the users to verify things so that the data team can stay focused on the top things that matter.

13:27 Uh >> so first you you had Jira ticket for give me data now you had Jira ticket for verify this data or verify this number >> right >> yeah it's the same >> it's the flavor has changed obviously it's downstream >> but it's it's still yeah >> yeah the intake kind of uh requests that comes through are different uh to be clear we kind of think of our data teams against an objective function we can talk about that separately so there's like bunch of work that the data team is irrespective of the intakes come through, but intakes that do come through do take up time and focus away from the main thing. And so we did wanted to kind of address that. And so we tried few different things. Uh we uh one of the things that we did is we made dashboards super easy to find and we narrowed the number of dashboards that are trusted that people need to go to. So on our platform which is what we use Quicksite. Quicksite is not great at being able to tell you which are like your top dashboards that we should trust. So we literally came up with an embedded dashboard concept in our internal analytics portal to show only the top dashboard so people know those are the ones that data team trusts. We literally renamed all the dashboards which were top trusted dashboards so that people can search for that keyword and then >> top top trusted or something like that.

14:59 >> Yeah, we need the blessed dashboards as a way to >> yeah blessed. Okay, good. >> The word doesn't matter the what matters is a unique name that no one else is using. So it's like easily searching >> and we made sure that it's like >> pretty small set of dashboards because if you have too many dashboard then it's going to defeat the b. You also made a domain specific. So like depending on the domain that le dashboard has all the context that that person might need to independently verify at least some of those metrics are correct. Uh that helped with some of it but it wasn't enough because dashboards didn't have segmentation. It didn't have drill downs in some cases. And so the next set of things that we did was we literally created a verification loop that can be run by the end users where the end users can literally copy a screenshot of the dashboard that they are seeing that the AI we need to independently verify. And as soon as they do that, then the internal plug-in that we use goes back and searchs for the dashboard code and matches the code that it wrote with the code that dashboard is using to kind of triple check that that's the right way that I've kind of started this numbers and then if it matches then it knows that all the segmentations beyond that are also correct. So those are those two things which is enabling the users to not only get the data but enabling the users and helping them build the confidence on their own that their data is correct goes a long way and they don't need to do it all the time they do couple of reps and most people would get a sense for like okay I kind of also have now the sense of the numbers so it's after a few reps most of our power users also don't need to verify because they kind of figured out and sometimes they just continue from where they left off so a lot of our power users actually ended up building their own skills uh skills >> they would kind of share with their kind of domains and say hey if you are in this and you are working on this install the skill and as soon as you do that then the other folks would not have to go through the verification uh issues as well. So that the learning here was mostly that once we got the adoption um it increased the work for data team not decreased it because we kind of had to help with more decisions.

17:11 classic paradox where I go back to this 2010s when Celsus BI came in people like if you have Celsius BI then data teams would be kind of less in nature but what has happened is that the role of data has evolved we have a family of data science which is doing higher value work very similar things are happening with the advent of AI as well where the value has kind of shifted to more high value and then but the unit of work has not changed it's actually increased So paras and maybe this is question also for you Shanti based on what you see I um I mean ultimately uh whatever we do from a it's a stoastic probabilistic model whatever we do it'll there still be random times there'll be drift there'll be random times it's giving you five correct answer I've seen this live five correct answer six line is off completely off Why? You just gave me five rows correctly and the sixth is completely off. Right? So that'll continue to happen.

18:14 How do we how are you thinking about one of the challenges I was seeing was like when you when it gives you false confidence because the five times it is answered correctly then you're like thinking okay as a human I don't need to look again or check again it's correct but we know it can go off the rails any time. uh how how are you guys thinking about it? >> Yeah, there there's a few things that we do in some of our projects to make this to mitigate the problem as much as possible, right? You're never going to 100% prevent it, but I can get to like 99.99 or something of that nature. So, we do a few things. We prompt in such a way that we give the LLM's outs if they're not certain. We want them to say, "I don't know." We want them to give up. Then there's the eval and harnessing piece of it which is the more interesting piece from a technical side. We build evaluations that run prior to data being surfaced to users. So user asks a question. We're not going to just run the question. We take that question, let the thing let the process do its thing, then evaluate it before sending it back.

19:28 Some of the things we're looking for are can we tie this back to a specific number? Can we tie that number back to the query that ran it? And then what was the process that ran after the query brought back the data? Does that make sense? For some of these, we build multiple metrics that we want to measure before we surface that finding. It's all dependent on how critical the process is, how much of this you want to put in.

19:53 Sometimes we'll have it go back and kind of be self-healing where you have it measure against those evaluation metrics, sometimes up to six of them, score all of them. If any of those scores is too low, say go back and do it again. You were low in this one. And we can repeat that process automatically until we get to a baseline score and then we can go surface that to a user. And so your so so your eval or your in a way I call it triangulation process is running independent like you come to one answer you come to another answer that was one part of what you said and the other one is plan before you do create your how are you going to do it and what's the success metric even before you do >> yeah so when we're combining those things together then right we've got our tool calls where each of the tools is basically executing specific code right that's non that's a very deterministic process it can only do one Uh, and then you combine that with now I've got a process that has to be able to link the output of that code to the specific result that it came out with to an insight that's generated. Great. We do those things and then we say go through this process, evaluate that against these metrics. Have we hit whatever that threshold is on those metrics? If not, go repeat some of those processes. Sometimes that means did you pick the right tool? Sometimes it means are you generating the insight correctly. So there's a couple things that you can evaluate there. But we found that process to be quite effective. Now, it's time consuming. So, you're not going to want to do this in like a live interaction where you need to be responding back in like 10 or so seconds, but for the type of dashboard or the type of numbers that you're going to be surfacing to a board to sea level executives, I'd rather be certain and feel comfortable than I would that it comes back quickly.

21:40 >> Yeah. Paris, how are you thinking about it? Yeah. Um, so beyond kind of the technical components Ashanti mentioned and I think all of them are super valid. Uh, the I'll give you an operator perspective. I think beyond the tech, it boils down to the culture at the company. So not all decisions need that level of certainty. So it might be okay for certain set of decisions and then at the end of the day the accountability is on the decision maker. So that is not shifted. Mhm.

22:13 >> A decision maker needs to be bought in that hey either they can work with a data science person for a day or two for the top things that they really matters where if we are wrong then we would be substantial kind of losses the business or we execute something that we can't walk back on. uh and so they are empowered with with DSS for the top things anyways but since the accountability is still on the decision maker when the answers come back at the end of questions that are reversible or they don't matter too much there are various gates that can be in place right you can launch a thing to only a subset of users and see what it does uh kind of staggered launch um lot of those kind of product launches we do are AB tested anyway so if we in case made mistake we can always reverse it.

23:07 Um, and so beyond the tech, there's all this cultural and process and people components that are equally important that at the end of the day, the accountability is still on the decision maker. As a tech and a data owner, my role is to build decision systems that can be trusted because trust is my currency and I'll do all the things I can to invest to make sure that tech is there. But with the tools that are in place, you'll never get to 100% deterministic to Shanti's point. I kind of make sure that my users decision makers know that accountability is important.

23:43 >> Makes sense. And I can add a little bit from our from Inola's perspective on our end. We are uh now launching with a lot of CFOs. We're essentially a command center for CFOs. So you know CFOs do the same thing over and over again every month, right? like they have to close the books and then they have to run their MR retention curve and then their P&L and then their their um you know uh QBR or MBR whatever those packages are they have to run like they know what happens week one of the month, week two of the month and so on. But one thing CFOs absolutely need is accuracy.

24:22 Uh because it's going out to investors or it's going out to um the the you know to the market or whatever whatever form it is. So it needs to be 100% accurate. So what we are doing here is kind of aligned to I think Shanti what you were talking about in Ola's main engine the the there are two parts the semantic layer and then there's the analytics engine. Both of them have significant huristics perspective like a deterministic perspective because we need it accurate. But for CFO where we can't even be off even a little bit we are we are almost like doing the 30 40 metrics because we know we've worked with CFOs for last 15 years. We're taking those 20 30 40 metrics and we are making it completely deterministic and even within the semantic layer itself.

25:15 So that even though we have automatic semantic layer that that central piece is completely deterministic so that cannot go off what what whatever may happen. So that's where it's 100% hallucination fee. That's how we are building trust from CFOs who can say okay I can trust this number. Ultimately the signer is still a CFO. They're going to sign on those packs, but they need in order for us to say your I don't know hundreds of manual hours saved because we are doing this all workflow all the way to the end >> for you to come in and just do 0.5 hours of where's my QBR? Here it is. You know, in order to do that, we really needed uh heristical method. Uh if we left it to the um I mean the other the LLM, it'll it'll not work that way. That is our approach as of now. Yes, to your point Shanti, there is latency is bigger for us. It's not we are not giving answers in 10 seconds. Our quick data answers are still I would say 25 30 seconds. We are optimizing it constantly and our longer deeper reports that get generated are still running two to three minutes >> but it's not something that needs to like people don't need to wait like it's a report that gets fired and then it'll get emailed to them. So it doesn't matter that two to three minutes doesn't hurt us that badly but yes we have to build a lot of these because what we do is we solve it one way and then we solve it another way completely independent >> y >> and that has to match so there's a lot of these things which we have done in order to make it I'm I'm calling it hallucination free because our core model is completely deterministic and using the sunundries on the side as uh so that's that's one way I'm also interested in talking about harnesses because there has been a lot of conversation. I don't I don't know whether we talked about it last time but I I feel like I was recently in a conference as talk and I heard CFOs says I never I never want to start a flow where I can't turn off the tap and they don't want to have commitment to one provider even though as of now maybe claude is winning over many other models for most of what we are trying to do. Uh so it seems like where the world is going is towards harnesses where people want to interact they don't want to they don't necessarily want to call these models uh individually or or have long-term contracts with them. They want to go via harness. I know what you guys are seeing around that or if there's a trend.

27:50 >> Yes. But then the model providers give you significant discounts to use their harnesses because they get more control over their infrastructure and how they use them. So like I'm not a huge fan of the clawed code UI, whether that's their CLI or the desktop app UI for it. Although they've added the new panels which are nice, but I use them both because I want access to the token discount I can get through a subscription. Same problem on the codec side. The codeex UI kind of frustrates me, but I want the discounted tokens. So, I end up using that they're harnessing and their UI to give them so OpenAI has control over which sub agents are going to be running and how that they how they handle those sub agents and how they cut off the memory to them and how they let them talk to each other. Uh so there's that piece of harnessing which is you know how it's going to handle some of those piece of the problem. Then there's the kind that we build for ourselves, whether that's around error checking or validation.

28:48 I find those to be a little more interesting because the user can choose how they want to implement important parts of their workflow, whether those are like research components, um, retry components, self-healing pieces, and you can build a lot of that into your harnessing. Or you can go to the open source community. I really like kilo code and open code and I believe kilo code rebuilt some of their CLI around open code. They kind of contribute back and forth to the projects. Uh and those are really nice harnesses for multi-agent parallel systems for coding.

29:24 And then you get outside of coding and in other other use cases and you probably want to look at different agentic harnessing. >> Yeah. And you and I have talked about it. I love perplexity because I can just choose to use 2.5 flash or depending on what I'm doing. I like the interface significantly. It's very expensive. I keep maxing out every seven days. I'm blowing through a lot of >> so something weird on perplexity. I've seen like what I thought were standard queries get routed to computer and I was like wait I don't want to computer this.

29:57 I just want regular this. >> Yeah. Yeah. Yeah. That just happened to me as well last I think this since last week. Yeah. >> Yeah. They they're doing something funny and I'm I'm I'm left wondering is it using more credits now? >> It's hard to tell because you never saw the credits on not computer, but for a computer you get to see all of the credits. I'm like, well, which what if I want more control over where this is going and how it's being used, but I get the sense it's all a computer behind the scenes. It's just a different set of agents and a different way of showing you what the agents are doing.

30:30 Yeah, I was thinking computer as more as an execution like an agent. You do get stuff done. You open this browser, you go and file this ticket in the Microsoft for my for me to create a marketplace entry for Anola, whatever. >> And then the the other one, the browser is just me asking, hey, what is this rash on my skin? Do I need to be worried? You know, so I'm I'm separating it that way in my mind. I don't know if that's how it's it's it's doing, but I am worried about credits because I'm >> I haven't looked deeper into it.

31:03 >> Yeah, >> like I I have a bunch of agentic browsers. I wrote about them, but I don't actually use the agents very frequently and the other day we were trying to debug an AWS issue and like why we couldn't get traffic to egress from an EC2 instance and we're going through the settings, we can't figure out what's wrong. I'm like, why don't we ask the agent? like we've got the AWS console open and a web UI. It can navigate that. Let's see what happens.

31:30 And like 10 minutes later, it was like, well, actually, your problem wasn't egress. The traffic was getting out. It was being blocked on the way back. So, we didn't know it was getting out because we couldn't receive the ping to know that it had come back in. It's like, just open that back up. Like, okay. I thought it was open. I saw accepting traffic from 0.0.0.0, but that's what it that's what it was. I was like, do you want me to do that for you? Of course I want you to do that for me. Why would I be asking if I wanted to click that button myself?

31:59 >> But good. Your brow your browser worked. I mean your browser. >> So that was uh that was Comet. So another Perplexity product. They were they're in agent browser or in browser agent worked great there. >> Yeah, I sit in Comet these days all the time. That's what that's probably why I'm burning through all my because everything I'm like I don't feel like doing this. Let me just write a prompt. Anyway, we digress. As far as what what what how do you thinking about harnesses and as a company I'm sure you're having chats with your CFO is thinking what's happening to the tokens token usage infra compute I don't know if you're having those conversations with the CFO or or the finance team and then how are you managing that and how are you thinking about you know obviously you're committed to snowflake so you are infra is committed you it's not like you have AWS and every other thing right like you're you're you're committing to one infra what are you doing with your tokens and so on how you're thinking about it.

32:53 >> Uh so a couple of things I think one tip I have for all the practitioners out there uh is kind of have a principled system architecture. So like in 2015 one of the decisions I made for the modern data tech stack is that for every part of the stack I'll pick the best tools available and create my architecture around it. So even when Microsoft and Oracles of the world were pushing for hey our platform can do everything I deliberately wherever I'd worked at picked the best tools for the job. So I was very early picking Lucer for example when they were at series A because I thought they kind of best at what they did and I picked like Snowflake when they were early. I picked tools that are best at what they did and created my architecture around it, not succumb to whatever the platform players were selling. This might sound obvious today because most of the data modern data tech stack is set up this way. You have DBT and your Snowflake and your kind of front-end tool as kind of fairly common architecture, but it was not that common back in 2015.

33:59 And 10 years later, we are at the same point where we have all of this like platform players that would want you to >> to use native. Yeah. >> Lock in to their ecosystem. Harness is nothing but whatever the model doesn't do. Everything else is basically harness. Your skills, your plugins, your kind of workflows, all of that is basically defined as harness. Right? So the architecture that I'm working backwards from is what are the various components that will deliver AI analytics and how do I break it and pick the best at any point in time and create an architecture so that I can swap in and out as needed.

34:37 um even within claude some of our routing is such that some basic stuff is handled by sonnet and the reasoning is done by opus >> and so we have created that's essentially a fancy way to say we build a harness where the model selection happens automatically on behalf of the user the user does not need to know what's using in the back end this is actually something claude makes it available out of the box too. There's this model >> that you can pick called opus plan which uses reasoning for opus and then tool calling and other kind of extensive stuff for sonnet which kind of reduces >> but we have created internal kind of rules and guardrails that would try to minimize the number of traversing that it has to do >> number of like net new contacts it has to kind of pull in especially when we know that we can route the agent to the right thing that it needs to at any given that like basically reduces the top token cost for us and so that's the architecture that I work backwards from and Uh it's very exciting time because I was playing around with GPD 5.5 it has kind of surprised me in terms of the data analysis and the data science capabilities that comes out of the box.

35:49 All of this models at some point of time will catch up to each other. I can totally see a world where companies can even have their local models if they don't want to send this data to third party. I can also imagine more regulated in industries and even like sovereign states basically saying that hey you can't use cloud API or open AI APIs you have to kind of do local stuff. So there'll be sovereign models and stuff like that for regulated industries as well. I can but the thing that is kind of useful for the practitioners is figure out your skills and learn the architecture where you pick the best thing for the job at hand and create your architecture around it and things will swap in and out and even faster than what happened with modern data tax stack and that's essentially the name of the game at this point.

36:34 Yeah, that's that's absolutely true in terms of swapping out things capabilities are growing, things are changing uh so rapidly. Um even if you were to take the example of I'm writing I'm in comet and it routes to computer for simple thing probably leaving a squeeze in my mind like is it using my computer credits? Is it costing me more? >> Yeah, things are things are of that nature. What about SLM? You guys any are you guys using SLM at all? We tested a whole Inola we with Inola we tested with a whole lot of SLM for some of the subm modules. It's 200 agents. Some of them we tested for particular reason for with SLM but haven't found it to beat uh some of the bigger models yet. Uh what about you guys? Are you playing with SLMs at all?

37:23 >> I don't know how how S is S in your mind. How small? Like I >> define SLM first. Yeah, like I was I'm running a 35 billion parameter model locally so that when I'm on the plane with bad Wi-Fi, I can still get some things done and that worked reasonably well. So that's the new Quinn 3.6 A3 30 yeah 35 billion I think is where that one is sitting. Um so we've done that relatively recently. Pretty happy with the results. I would maybe call that medium sized and not quite small. We fine-tuned a 4 billion parameter Quen 35 recently and benchmarked that against the base model and against Sonnet for a very specific use case and we were able to see the improvement we were looking for across both of those. So we took Sonnet which you know that's like a trillion parameters with some custom prompting to say act like this and we were still able to outperform that with a fine-tuned four billion parameter model. So pretty happy with that. you know, competing with folks a thousand times your size or so.

38:25 >> Yeah, that's pretty interesting. And the latency probably I mean, is it beating in latency as well? >> Oh, yeah. I fast. >> Yeah. >> So, this one we were actually running on a runpod GPU. So, we were running it with like 96 gigs of VRAM, which is plenty for a 4 billion parameter model. >> Yeah. Okay. So, it was not local, but it was still faster. >> Yeah. Local depending on your hardware and things, you may or may not see that type of inference improvement.

38:50 There's a lot of dependencies there. There's like one of my machines I could run that on locally and it would probably go pretty quickly. I was running this one on run pod because I believe that one was in the normal safe tensor format. So, we're just running it as is. If I if I was running it on like my laptop, it would be in MLX and that will make it slightly slower than running it as safe tensor.

39:16 >> Yeah, >> I need I'm I'm I have a Mac. I need a Linux machine. Yeah, you you are you are you on Linux, Chanty or? >> So, I've got a desktop with a triple GPU setup that is a Linux machine so I can run CUDA and GGUF style models there or if they're small enough like standard safe tensors and then I've got a MacBook Pro as a laptop so I run MLX models there. >> Got it. Paris, how are you thinking? Are you guys using SLM at all or >> other teams at open door have played around with it and they found use cases uh where SLM kind of worked for what they were trying to do on the data side?

39:52 It's on my backlog. >> Do you know which ones? Do you know which ones? >> I'm always intrigued with what people are trying. >> No, I I don't know the tech. I need to It's on my backlog to try this out. Primarily because right now the tokens for all of these models are so subsidized that the opportunity cost to just focus on the key thing and figure out the architecture beats the uh ROI to even try the small models at this point. At some point we will get to it once these companies go public and it's not subsidized anymore. But uh at least right now I'm focused on learning about the harness architecture and that's the highest value ROI compared to switching models at this point.

40:37 >> Yeah. Okay. The last thing I want to talk about I know we are at the top of the hour and our our time together but I want to talk about this new concept people keep I mean it's likely the creation of a marketers from a c from market marketers naming another thing another thing uh I heard human human aentics architecture humanent did you hear that human agentic architecture hasn't reached you okay I just read >> I I don't know that one >> it just looked like it was saying humans need to be in the loop or arguments need to work with the agent. But there was a new name, so I had to be on top of it.

41:12 Who is this that I don't know yet? It's probably uh uh it's probably marketers doing their job naming yet another thing for an obvious thing that I think I think the future has humans have to be in the loop. They can't quite I mean I think it's an antithesis to the autonomous the the whole idea of headless and autonomous. This is what it looked like to me, which is like we don't have to touch it at all. It'll come and do stuff and here I'm the CEO running a billion dollar company with 100 agents or something like that. You know that story >> is always going around. I think it was an antithesis to that saying no humans have to be in the loop and humanentic is the new architecture to think about in future. I don't know. Paris you you were saying something. I think it to me it boils down to what problem you're solving and what workflows you're trying to automate. So at open door we buy and sell homes algorithmically and humans are in the loop because the unit of transaction that we do is at least a quarter of million dollars at stake. So increased the number of human in the loop significantly but it's not zero. Uh so there are parts of the processes that we able to automate end to end and lean on AI and traditional ML and have been on this journey for the past like 10 years and I've been part of the journey for the past four years. Uh but to me it boils down to mapping all of the various workflows knowing kind of what the downside is if things go wrong.

42:48 uh boils down to what not your what your median performance is boil down to what your tail performance is because if your tail performance is bad then you essentially lose bunch of money. Uh >> wait what do you what do you mean par what do you mean tail performance? >> So your 5% of the last kind of product would incur 95% of the losses. So if you cut those 5% of the products then you kind of are profitable right? Uh so being able to focus on those like tail outcomes and making sure your systems is able to flag those tail outcomes before they happen. For those things, humans are better off being in the loop. And at the end of the day, it boils down to map out like your all of your decisions that you're taking, map out the worst case that could happen, map it to your P&L and see what your kind of risk appetite is. At the end of the day, it boils down to that as well. Speed versus risk in this case. Uh so you can go speedier by giving it to agents but if your risk is too high and you can't like stomach that risk then you don't have a choice but to put a verification loop. Now whether that's another agent or a human it's up to you to kind of design the system. So at the end of the day trade-offs and high level kind of judgment trade-offs you still need to make those trade-offs uh that matter.

44:03 I was losing you a little bit but I think in summary what you're saying is you know map out your process and then look at high value where your risk is high you definitely have humans in the loop and then where your risk is low uh see what's at stake and think about how many humans can be swapped out from the process so you're talking about a hybrid hybrid model is what I >> well cool I am going to stop recording >> thanks for listening to not another AI podcast. If you found this conversation useful, consider subscribing so you don't miss future episodes where we break down what's actually happening in AI market beyond the headlines.

44:45 See you in the next episode.

Summary

Paris, the head of data at Open Door, discusses the transformation of the company's data practices over the past four years, moving from a decentralized structure to a centralized data team. This shift has led to a significant reduction in the number of data personnel and a transition from traditional dashboards to AI tools for information retrieval.

- Open Door initially had a fragmented data structure with multiple analytics tech stacks, leading to a lack of trust in data.
- The company consolidated its data teams from 80 to 25, focusing on higher talent density rather than quantity.
- A shift occurred where 80% of the company now uses AI tools for information instead of dashboards, marking a significant AI transformation.
- The implementation of a semantic layer and data catalog has been crucial for ensuring consistent data across various AI tools.
- Open Door employs a hub-and-spoke model, allowing data teams to partner closely with business units for key decisions.
- The verification process for AI-generated data has evolved to empower users to independently verify metrics, increasing user confidence.
- Paris emphasizes the importance of balancing speed and risk in decision-making, advocating for a hybrid model where humans remain involved in high-risk scenarios.
- The conversation also touches on the future of AI tools and the importance of flexible architecture to adapt to rapidly changing technologies.
© transcribe · For agents Built with care and craft by Gokul Rajaram