Section Insights
Implementing a Chatbot
How do I go about investing in a chatbot for my overwhelmed team?
To successfully implement a chatbot, it's crucial to have your underlying data organized and accessible. The AI needs to work with well-structured data to function effectively.
- Organize and structure your data for effective chatbot implementation.
- Investing in a chatbot requires understanding the underlying data needs.
- A well-functioning AI system depends on the quality of the data it interacts with.
Understanding Context in AI
How does AI understand context in user queries?
AI systems can interpret context by recognizing relevant information in user queries, allowing them to ask clarifying questions without needing all details upfront.
- AI can infer context from user input to enhance interaction.
- Contextual understanding allows AI to ask more relevant follow-up questions.
- Effective AI communication relies on understanding user intent and context.
Limitations of AI Models
What are the limitations of AI models in handling inconsistencies?
AI models can identify inconsistencies but cannot resolve them without human intervention. They require guidance on how to handle missing or ambiguous data.
- AI models need human oversight to address inconsistencies in data.
- While AI can detect issues, it cannot autonomously decide how to fix them.
- Human involvement is essential for effective data management in AI systems.
Evaluating AI Responses
How do we assess the accuracy of AI-generated information?
AI responses can be categorized into four buckets: absolutely correct, absolutely wrong, partially correct, and partially wrong. This helps in evaluating the reliability of the information provided.
- Categorizing AI responses aids in understanding their accuracy.
- A structured evaluation approach can improve trust in AI outputs.
- Recognizing the nuances in AI responses is crucial for effective use.
First Impressions of AI Technology
How can we ensure a positive first impression of AI technology?
To avoid negative perceptions, it's important to thoroughly test AI systems internally before public release. Understanding where the system fails is key to improving its effectiveness.
- Internal testing is vital for refining AI systems before launch.
- Identifying failure points helps in enhancing user experience.
- First impressions matter; ensure AI systems are reliable before deployment.
Transcript
0:00 Yeah.
0:00 So Lohit, you've deployed many, many chatbots, conversation AI products in the past. So let's say I run a team and we are overwhelmed by ad hoc questions, just random questions left and right about our business and we really want to invest in a chatbot. How do I go about doing that? What needs to get right in order for something like this to actually be possible?
0:26 I think that that's a, that's a great question. I think a lot of startups should be already these stuff. My, my understanding and how I probably will build this system. So, so I think the most important part is see it being something like
0:40 that with like AI feature that can be hooked up. You have to have all of your underlying data and stuff and implemented in a certain way for the AI to actually work well with it. That's probably entire job in itself.
0:54 Exactly, exactly. Because just think about this, how dangerous it can be. So you go, you go with the data set which is comparing quarter two and sales data and I don't know because of hallucination, if the number came opposite, you go and present the same portion to the executor, guaranteed.
1:13 I guarantee there are going to be some stuff where it's like someone did their earnings report publicly, they did that, the AI got it all fucked up, their stock went crazy and then they have to come out and just like, oh man, like false report, CFO's fired. Like there's going to be shit like that.
1:36 On the podcast today we have Lohit Yogi. He's a technical principal product manager at ServiceNow building Agentix Systems and has spent time building AI agents at Adobe. In this episode we talked about chatbots versus conversational AI. See the LM revolution firsthand. Natural language business intelligence. Are AI agents ready to answer your data questions? Real world AI deployment challenges, hallucinations, inconsistencies and data governance. Building AI data agents, strategies for success from UX to managing errors.
2:11 Be sure to subscribe to Data Neighbor in your favorite podcast app and YouTube. Let's dive into our conversation with Lohi Jogi.
2:20 Hey everybody. Welcome to another episode of the Data Neighbor podcast with hi Shravya and Sean. Today we have Lohit Yogi with us. Welcome Lohit.
2:30 Welcome Lohit.
2:31 Hey, long time and yeah, certainly we're going to have a really good conversation with you about conversational AI. So before we jump into our topic, you've been involved in the whole AI space, specifically on the product side since even before the ChatGPT moment came out. What's can you tell the audience Kind of like your journey and how you got into this whole thing from the very beginning.
2:58 Yeah, sure. And then that's a great question. So. And thanks, thanks for inviting me to share my insight what I learned with my career. So like my journey into a product leadership started just after my school. So when I finished my school I started working as a machine learning engineer. Fortunately I got a lot of opportunity where I can showcase my technical skills and leadership skills. I started with a very like mid sized company like Call Asura.
3:25 It was like 2018. So like that time we were building chatbot for the insurance company. So that particular experience taught me the importance of like how do we understand the customer? How do we use that knowledge that okay, why customer is coming to us, how do we help them? How do we understand their problems and provide them the right resources. So like just think about that at that time. We started with like very simple basic regex models, moved from regex to machine learning models models, then deep learning models.
3:57 And the main idea was that okay, how do we help them? I worked there for a year or something and then I moved to Adobe. I stayed In Adobe like 5 years, similar type of things. So I started with a single text based conversation AI where you enter a text, we understand the text, use that text to identify the intent and give you right resources. We successfully did a lot of good things and then we slowly made it omnichannel.
4:25 So if you call us we will also understand you, provide you the right resources. It's very hard on the phone. Nice for everything. Yeah. And then we move from that deep learning transformer type of model to a large language type of models. Right now I'm working with ServiceNow. We are very, very dedicated to the how do we develop agentic workflows and then automate this. So the main important part where like I move from after my school to a very specific product thinking is when you start using like I'm very, very passionate about machine learning.
4:56 So anyhow I wanted to use machine learning for the real people when I started my career and when I worked first time in a company where they were using machine learning to help customer, that was amazing. So, so just think about this. When you have a limited resources as like a small size company and someone is trying to reach as a consumer, like I'm a, I'm a consumer, I'm trying to reach out to you. Can you guys help me with this?
5:19 And if I don't have enough human support then all of those customer will be waiting in a line. Yeah, like A basic teller example. When you use machine learning in all of those scenario to reduce that overhead. That's amazing. Like amazing. Like you feel so awesome here. Like I don't know if you guys also go into bank and if there's a one teller and there's a big line, you will be like. And buddy, if you have an app which can do the simple things, you're good.
5:46 I only go into the bank once every four or five years now though, I feel it's like it's all digital and now I. The last couple of times I've gone in there it's been like I have to like, I get like a check from like another currency and I have to go to the bank to get it deposited and I'm always like, how do I do every everything exactly right.
6:06 So just think about that. Even that's a one basic example. So the bank apps used to like, they don't allow, they used to not allow to deposit a check using app. Now you can actually use the mobile app, take a screenshot of your check and then deposit it. You don't need to go bank at all if, if there is something, if you don't need cash. And all of this is happening like sometime with the business intelligence because you are taking a screenshot, they are verifying the screenshot, they are doing some backend system to verify that okay, the signature is correct.
6:36 That's definitely coming from like some ecosystem which compared both science and machine learning. And then say, oh yeah, this is the right person who are depositing the check. So we have got, let's deposit this. So, so like these are the basic examples. So my true like motivation was that okay, machine learning is awesome, like really, really awesome. And how do we really use this to build products, feature or different things which can really help the customer in the real world.
7:02 Yeah, no, that's awesome. So I picked up two key terms that you kind of alluded to in that answer that you just gave, which is chatbots and conversational AI. So what's the difference between the two is one, are they synonymous? How should the audience be thinking about them?
7:21 I think they are like pretty similar. Like it's just a different terms. So the conversation AI in general we call as a broader scope because like chatbot in general is developed because it was chat and there is a bot to help you in the chat or say the conversation AI is kind of like the overall modality conversations between voice, image, video, everything understand and like use it. So so I think this is the basic difference. But like it's a very interchangeable like people should be able to understand both if you're using an as a vocabulary.
7:51 Okay, that's fair. So this, this, this may be an obvious question to you, but hopefully you can paint a picture. What are the limitations early on with chatbots that you build with machine learning models? Ver in terms of the large language
8:05 model powered ones, I think that that's a great question. Like I'm going to like try to give some real world example in this. So, so just think about that. The earlier stages, the chatbot was very, very specifics. Like they, they, they were very limited in terms of understanding. They were very limited about that. Okay. You have to write in a certain fashion to get the right response. But if you miss that, that what you're writing, even misspelling can make it, make it bizarre.
8:35 Right.
8:36 So for example, even just like you say what's like your return policy and if you misspell return in some fashion which can translate some other word or its closest match to other word, it will automatically start like it will not understand you. Yeah, it has to be like, oh, I don't know what you're saying. Can you just say it again?
8:53 Yeah, especially like the phone ones when you're just like yelling at them to trying to pronounce the word correctly. Like I know I said cancel subscription human.
9:06 Exactly. Right? So, and then it's like, and, and this is very frustrating. Right? So it's not like a simple, that, okay, we are just repeating. But they had a like limited understanding about the context. They were very, very like basic to just like matching what you're saying. Use that information to like find if we can, if we cannot just like say something. The second one I think I failed is like they were also very rigid, right?
9:28 So rigid means is kind of like they never maintained any context. Like when you like you just made an example that you were saying cancel. And then what if you say my plan, they will take both of them in a different as an input. Right? They will think cancel was something else. Now my plan is something else and it will automatically give you some weird result about your, my plan. And it will generate like a lot of like frustration for the customers because like what you're doing is that you're not keeping any context.
10:02 You are just like thinking that every word is coming like a brand new question. And then you use that information to find the relevant resources, which is totally wrong because no one talks like this. Like, they don't have a tendency to just say everything in One line. So sometimes people pause, sometimes people just say okay, yeah like it's kind of like they think and they type and which actually break all of these older type of ecosystems or conversation or like chatbots.
10:29 A third thing what I feel like is like more or robotics, right? So before the generate your models, everything was handwritten. You find an intent and then that intent supposed to trigger a workflow which is defined by us or the people or the engineers and everything. So exactly one on one mapping. If you take this decision, then only we take could go to the this decisions. Very, very robotic. Very, very predefined notions. Very, very like there was no intelligence behind it.
10:57 This was just like okay, they are saying cancel show this line. They are saying buy show this line. They are saying billing show this line. That's, that's like, like a Magic 8 ball, right? So you ask the question. They have like some repeated number of things and they will just keep on showing that. So it was like pretty, pretty bad. And now if you compare all of this like older decision tree level or regex level chatbots to another current ecosystem, you will be so fascinated to just see like it happens.
11:26 All of this happens in like last four years. So the rule base is gone. Now we have like large language model which is not just understanding the customer intent, it's also generating the responses what customer wants to see. They're like, we are also keeping the context of the overall customer conversation. We don't need to actually just like think about oh, everything is happening in a behind the scene and like handling ambiguities, handling the multi turn conversations, taking action behalf of the customers and then how they're thinking what they're thinking.
11:55 Like a basic example is that if I go to a chat GPD today and then say I'm planning a trip to Seattle next weekend. So just think about that now. The system is so intelligent that they next weekend is also a context for them. So they will not ask you the dates. They may ask you okay, are you going on Friday and coming back on Sunday? Is that you're saying or are you growing on Saturday or coming back on Sunday?
12:19 Because they already understand that the next weekend is, it's a context. It's not just about that. They are just saying some random information. So like this is pretty, pretty incredible. I don't know. Like if you guys want I can also give you like some context that how what's happening behind the scene. So like when you ask these type of question assistant understand the context. They try to build this relation between Your query and the context.
12:44 Okay.
12:45 And then they try to see that if there is ambiguity, what we are doing. So for example, what I said is that next week kind of like or weekend is kind of like a specific date range. So. And then I'm saying that I want to book a trip. So my intent is really that, okay, I want to book a flight or something. So asking something that, okay, yes, I can find you a flight or car or anything, like, different options, how they want to travel and giving that option back to the customer will help them.
13:14 Okay, yeah, that makes sense. That, okay, I am doing this and this is the way I'm adapting all of this conversation and then building the conversation according to the customer inputs. Instead of like a predefined notions that you click on car or you can click on flight, then go to the flight mode, it go to the hotel mode and stuff like this. So. Which is like, quite amazing to see that all of this has happened in very short time of period.
13:38 Oh, yeah. I feel like I've gotten so used to just like anytime I'm working with like an AI, whether if it's in like a chat bot, or if I'm like writing code through some sort of like, agent coding tool, or if I'm just like chatting with like, ChatGPT or something, it's like I now just like brain dump everything. I'm like stream of consciousness, like, full of spelling errors. And then it's like, okay, yeah, I think this is what you're trying to get at.
14:08 Like, I could send that to my wife and she will not know what I'm talking about. She'd be like, I don't know how to interpret this. We need the AI to interpret this. But which is so crazy. Like you were saying before. Whereas before, if you just like, misspelled a couple of words, it was like, I don't know what you're trying to ask, when to ask it again. And now it's like, hey, I think you're. Yeah, like you just said, you're trying to plan a trip.
14:34 Here's all the workflows we could do together. Like, clarify. It's pretty cool.
14:38 Exactly, exactly. And then being intelligent is exactly like this. Right. Even small things. I don't know. Like, most of the time I used to do all the maths with my hands. Now a lot of things, you can just like, put it there.
14:51 Yeah.
14:51 And like, you can just say, oh, 20% of this or 50% of this, or calculate the 90th percentile of this. Or like, things like this like it's pretty, pretty awesome what we have right now. It's very, very powerful.
15:04 Yeah, yeah, no, for sure. That's, that's really cool. Hey, I think we're going to get into this really fun section here, this really fun segment here, which is you have mentioned you're really passionate about building chatbots or you've been very passionate about having developing chatbots to help customers. And obviously on this podcast our. A lot of the customers here are data professionals. So whether they, they're. So let's pick maybe like one segment of them. For example like data scientists, they do a lot of, they get a lot of questions throughout their day.
15:36 Like hey what you know, like can you plot me this? Can you pull me this number? Can you do X? Can you do Y? That would seem to us to be a very perfect situation to use some sort of chatbots to, to help resolve those sort of questions. So a question in this vein for you, look, which is how close are we to natural language business intelligence where we can just ask questions to a chatbot and then it will be able to provide us with answers according to our databases, our own data and it knows what to do and what to compile.
16:14 Yep, I think that that's a great question and my, my response on this is that we are almost there. Just think about this. The overall conversation AI is evolving so rapidly and it's just not like helping about the data scientists, data analysts or non technical stakeholders.
16:32 Yeah.
16:32 How do we actually put all of that information in a fashion that can be like leveraged using the conversation AI. It's kind of like that integrating both of these ecosystems and if you see in the back end side, no SQL or the graph SQLs. Right. What they're doing actually is they're using all of this data tables in. In such fashion and building those intelligence around it like those, those different like older SQL was like very limited.
16:59 They have tables and table have like maintained their own key value pairs and all of those things and unique value enable a thing. Now the GraphQL is not just like basically just maintaining those ecosystem and tables. They're also actually building the relationship between the entity and nodes and then like how do we drivers the nodes and entities. So like just think about this like this ecosystem when connect with the ecosystem where you're asking things in the natural language processing or like a simple text or voice anyways it will take that query, translate that query with a very similar or generate SQL query and then find out what should be the right things they need to Fetch out of these tables, use that information to summarize and build an insight.
17:46 That's pretty awesome. I know that. A lot of challenges behind it.
17:49 Yeah.
17:50 But just think about, this is like a brand new junior analyst which can make mistakes, but still it's kind of like can do the 70% of job done. So you need a supervision on top of it. Definitely. But it's kind of like a junior analyst which can write a simple SQL query for you, find out the right information, build that information behind the scene, and then you can just ask him, can you give me the total sales revenue between X year to Y?
18:15 And that should be able to give you the information. Again, a lot of challenges behind the scene, but still that's doable in some fashion with some accuracy.
18:23 Yeah, I agree. I think we're getting pretty close. I mean we had. I remember it must have been like two years ago or something. I don't know if you remember this high. We're at next door and we had the like, like head of sales for OpenAI came in and gave a talk to like all the data people. And I remember at one point she was like showing off open AIs. This was like, I don't know if it was.
18:47 Maybe it was three years ago. It was like a year after like 3.5 came out, maybe even six months after. Not, it was probably six months. And she was like showing off like, oh, and you can upload an X CSV and like now it's like having a data scientist in your back pocket. And I remember being like, hey, first you gotta know your audience. Like you're just telling us that we're gonna lose our jobs. But second of all, I was like, no, I've used it like it sucks.
19:15 But now when I'm building out this evaluation pipeline at work, and a lot of the way I'm building it is I'm using cursor and I'm using clause code. And so it's actually like it understands a lot of the pipeline itself, like you were saying, and understands a lot of the relationships and the tables. It's actually building and writing to Snowflake on its own. So later on when I'm asking it, okay, now help me build a streamlit app that like analyzes these results in this specific way.
19:46 By the way, go look at like the like build table function that you wrote before to understand exactly all the input data into it. I can actually do a pretty good job because it understands all the context of what's how, how it's been built behind it and some of the business logic. I think what I'll be interesting to see though going forward is like that business logic piece for like all these established like data warehouses and stuff where someone's trying to like plug an agent into it, but it might not have all the context of like what everything means.
20:21 Yeah. I don't know if you have any advice in terms of like for people who are trying to annotate or what they got to do to basically make this stuff as easy as possible for AI to understand.
20:34 Yes, I think that that's a great point. And like so my suggestion is that like it's not about losing job. I think the most important part this advancement is that how do we use AI to fasten the stuff and do more like effectively. Like that's the way I always see. So just think about that. Like in three years ago if someone asked me go and find out the key trends in the sales report that may take like a week, two week or three week depend upon so many factors that what are the different types of joints you need to do, different types of exploration you need to do?
21:09 What are the different tables they are using? What are the, what are the right value you need to use and all of those things. Right? So finding out like these type of things. Oh, why like next question that once you have like okay, key trend and now the next question is with the leadership ask in the same meeting where you're presenting what are the revenue drop in Q2 why that happened. So that's kind of like generating a new question and then takes like one more month to just come back and then it's like so, so slow.
21:34 Right? So it's not about losing the job, it's just more about that how do we use this now to get the better insight faster so we are not hung up on just like just doing some data stitching to find out the right information. Right? So it's about this, right? And think about this. Even like in all of those times I had to clean the data, transform the data and then if something is missing in the column that could just like add the noise and so many normalization other type of like exploratory analysis on top of all of those things can just change the result dramatically.
22:10 Because like sometimes you think that you are doing the right things and then you go and present and the leaders will say I'm looking this data from last three years, I don't see this. This is right. And that's going to happen. And new question is like totally different thing. So how do we use these advanced model to just like bring all the knowledge on the table. So when the executive is making those decision related to the product, they are like very confident.
22:33 Yeah, this looks seems fine and we should do this. But again, again this, all of this come with the challenges. Like I don't know, it's not very simple. It looks very dreamy. Yeah, I can understand that. Oh, when we are selling these things where we are asking, oh, it's amazing. Looks pretty simple. But it's not very simple because there is no one solution on this. Right. For example, just like basic one challenge. I feel like there are a couple of different channels.
22:58 But yeah, one channel is my system is working with the 85% accuracy or 95% accuracy or maybe 99% accuracy still there can be schema ambiguities. Like for example, for me, Churn could be different versus you could be using Chan in a different fashion.
23:12 Yeah, totally.
23:14 That table name you are using the table field or the column name you're using could be totally different. Right. They can be like okay, what's my active user? An active user can be distributed in different segments, right? Daily, weekly, monthly and blah blah blah. An active user just who log in. What's the definition you're using with your table? Still relies on the people like data engineer, data analysts to own those metadatas. They're just like using this tool to fasten this process.
23:39 But this ambiguity cannot be resolved without a real human because machine will not know what revenue defined according to you. If it's a quarterly numbers, it's a weekly number, is a monthly number. Like what are the different things? Like you still need to define all of those things. And the other things is like inconsistencies. Right. So these models will not fix the inconsistencies. They may remove them, but they need to specifically tell them okay, if you find something is missing, fill it with the average.
24:07 Right. Like still a human needs to say this. They won't be able to take this decision without a real human at this moment. I don't know about the future. Again, like things are moving very fast. But these type of things still require someone to understand what makes a lot more sense when we are training the data building these models using these tables. So like all of these things and I don't know like so these are the couple of things what I remember which we see as an inconsistency.
24:34 Yeah.
24:35 Even though all of this still LLM can be so confidently say wrong things.
24:40 I know, yeah. The humans you're safe for. You're safe for a while, it seems like. Yeah. Reiterate what you said. I totally agree with you. Like about like you're not going to lose your, your job kind of thing. It's more about like getting things fast done. Getting things done faster or getting things done that you just couldn't do before. Like every time I leave a company I'm like, hey, here's my like Excel or Google sheet of like the 200 things people asked me to do over the last three years.
25:08 I never got to because like there's not enough hours in the day. It's like now maybe we'll have time to do that. And I really like your response around like, yeah, what happens when you get like a question from an exec, go spend three weeks, try to figure it out, come back and they're like, that's not exactly what I was asking. And then you come back again. It's like, now you can do that so much exact professor.
25:30 And then. But I do. Yeah, I really like what you're saying. It makes a lot of sense in terms of like the like business logic and definitions. Like that's such a key part of like why data analysts, analytics engineers and data scientists exist is to like talk to all of the right people and people they partner with and teams to like come up with these definitions. That's like the heart that's like probably the hardest part of the job.
25:56 Every company I've worked at with, like, especially when you're working with stuff that's like the executive level is, is looking at or, or if you're at a public company and you're reporting stuff to the street and it sounds like we really have to like in the data profession like double down on that to make it even more transparent for an AI to be able to understand it beyond even just like ourselves and our stakeholders if we want this stuff to work.
26:21 And then the stuff that gets offloaded is like, yeah, the complex querying and joins and that kind of stuff. Like, but it has to know that business context in order to work.
26:31 Exactly. And then if you think about this like this also like give one more opportunity of data governance. Right?
26:37 Yeah.
26:38 Right now if you put everything inside the larger LLM models, what they will, they will analyze all of these things and then they will generate insight, but they will not think about that. Something sensitive and non sensitive. Right. Data governance, whole data governance, privacy and everything still needs to be defined. Someone who understand all of this. So they will not know okay, if what is pii Versus what non pii they can expose anything. Right. For example, recently we are working on a problem statement where actually we release something.
27:09 So the idea is that if I ask to my chatbot like who's my manager? It should be able to give me an answer that who's my manager? Because of course that's the public information between my org. But if I'm asking what's my manager's salary? The table is there.
27:23 Oh man, that's cool. I gotta start asking that stuff once it gets released way this is.
27:31 This is not like a. Like that's what I'm saying. Like this is not about like. Because like why would I lear would not answer this to you. But this is like a sensitive information and still the table has somewhere this information and can be fetched. But building data governance around it's very, very important because you cannot just expose anything. What you think about this and this like the overall like whole PII things and if I give the more smaller examples is that like you say okay, analyze this data like it's small table.
28:01 Just think about the very small table. You say, okay, take this data and build a insight for me and build data visualize some kind of like maybe bar chart or anything. If you don't tell them which type of chart you need to build, they will still be confused what they're talking about. So it's like very, very simple things. It's not like you need to. A person needs to know that okay, this type of data looks better if I build this type of charts.
28:28 Yeah, and you need to make them learn with the time that okay, like for example type time series, require line chart, category comparison, maybe require a bar chart like distribution require pie chart or histogram and stuff like this.
28:41 So.
28:41 So all of this is like a knowledge which actually depend upon that what type of analysis you're doing and how you want to present this data. And LLM is always confident even though they are wrong, they will find out a response which is like you like okay, maybe I am wrong. I need to figure that out.
28:59 Yeah, I have found that where you're sitting with the charts like if I'm building like a streamlab app or something, even if I give it like a prompt that I've had like another LLM tell me a bunch of like data visualization best practices. Like if it tries to one shot the app it's just like the color. One thing that always happens actually is the colors. It'll be like white text on like a light green. And you can't even read it or anything.
29:23 Like it's like human readability isn't quite there. And then just like you were saying like weird chart selections and it's kind of like you do still have to go through and be like pretty specific about like your design. One thing that is really cool though, with like cursor you can like upload images to it. And what I have done is I've sketched out what I want it to look like after it did like a first pass and I put some labels on it and then I took a picture of it and I was like, no, make it more like this.
29:51 And it still didn't get perfect, but it like got a lot better, which that was pretty cool. Like the multimodal interpretation and then like creating an app out of it.
30:02 I agree, I agree. And that's why I'm saying like so like thinking in a like a different level, it's very important. Like how do we use this for faster delivery, discovery? How do you use it for better data clean cleaning or implementation? How do we do a better like a quick edas, like exploratory data analysis? How do we find out the patterns, how do we find out reusabilities? Like all of those things is pretty important, right?
30:30 So like and then then like with the time of course, like we are learning all of this and then you will like all of this workflow will make it faster and faster which is reduce the effort. What you have like in the older time, you have to read everything to find out the insight. Just a basic example is that any chatbot, like for example, even dedicated company chatbot, has millions of conversation every year. Millions. If you ask me to go and do an analysis on all of those, I would be choosing like thousand maybe, right?
30:59 Like that's the only way I can actually use this data set to build a 90% style confidence that okay, what's happening with the customer, customer journeys. And that will not still represent the overall picture of those millions of conversations with these type of tools. Actually you can do all of this and build a like a better confidence at what's happening there. Again, it's not perfect still, it's give you garbage results. I'm not saying that you should not read the conversation.
31:25 You should always double check what's happening. Hallucination, is it real? It can confuse you. It's so confident that, okay, you don't need to always, always double check. And that's why I don't know if you see that a lot of ecosystems, they give you the query and the insight. So they give you a query which can, you can understand what the system did and why this insight was generated. Very, very useful. But that query also can be wrong.
31:50 So, like, still, you need to use the human supervision to understand what's happening there.
31:55 Yeah, no, that makes sense. So, Lohit, you've deployed many, many chatbots, Conversation AI products in the past. So let's say like, hey, I run a team and we are overwhelmed by ad hoc questions. Just like what we just talked about in terms of data requests and just random questions left and right about our business. And we really want to invest in a chatbot. We understand there's hallucinations and things like that, but we're like, hey, we, we, we got to kind of stop the bleeding, if you will, and have a chatbot help us out going forward at a high level.
32:33 How do I go about doing that? What needs to get right in order for something like this to be in place, to actually be possible?
32:41 Yep. So just to clarify that, what you're saying that you are getting a lot of requests about the data and you want to build a system which can take all of those ad hoc requests and respond back. So how do we build a system around it?
32:55 Yeah. So instead of getting the questions to the data people and they have to go figure out the answer, it's like, hey, the bot just helps us kind of shield all. They would just be instantly, instantly smart enough to help with 247 sort of questions around the clock.
33:15 I think that's a great question. I think a lot of startups should be already working on these stuff. So I'm going to give my understanding and how I probably will build this system. So I think the most important part is that there are a lot of pitfalls and there are a lot of challenges. But I'll go with the ideal scenario that there is table, which is awesome. They are building the very new technology and there is no latency.
33:40 And the best way is that you build a UI or Power BI dashboard, which is connected with a chat ecosystem where you can see some dashboards, you can see some information. So whenever end users start interacting, they also understand what's available behind the scene instead of just like just giving them only a single bar. Right. For example, ChatGPT gives you a bar where you can ask anything. That's like very, very open space to do everything. I don't think so.
34:07 That's a system we need for data. Ad hoc system. We need a system where you expose certain metrics in front of that chatbot. Or conversation, AI or BI conversation, whatever we call it. So the consumer at least gets some hint of information. They see some chart, they see some high number of conversation or number of question asked by customers, most frequently asked questions and all of those things. So they know that okay, this has already been asked or this has already been there or this type of information exists.
34:37 Like I'm not saying that expose everything, but some sometime up to our mission. It's very important. The second part is that why it's important because if you give them the open ground, they will ask so complex query and then you will just break the system in very, very simply.
34:50 Right?
34:51 So just think like this. That already defined ecosystem, some knowledge about the existing KPIs testing joins existing data where you can fetch the information directly. So all of this showing this is very, very helpful. The second thing I think probably I will think about this. Like as you said that a lot of customers coming there and asking ad hoc queries. So how do we provide a chat bar where customer start asking questions and how do we translate those questions into a query and then that query to fetch the data from the backend tables.
35:29 Yeah, this is the way I will think. Now the part is that when you are fetching so converting text into SQL looks like moderate to difficult problem. Fetching data with the XQL query is also like a moderate, moderate issue. Fetching the right information, that's kind of like a difficult problem to solve. So converting SQL maybe SQL query is right or wrong or maybe partially right, partially wrong that still require a human like intervening like okay, how do we.
36:00 And again, just think about this. We are building a new system. So we are also thinking about that. Okay, with the time, how do we improve this? So thinking about this in like a four bucket that okay, it was absolutely correct. What's it absolutely wrong versus partially correct worth it? Partially wrong. And this is the cue we add when we fetch the data and show to the real time user versus that. How do we also define that?
36:21 Okay, when we started fetching the information, where was the occurred? So for example I used a field which does not exist in the database at all. For example, if I say give me information about distance or maybe something like sales numbers from a product which is not your org. Okay, give me the number of Uber. Like that's like you are working in a different company and you're asking numbers from a different company and maybe information doesn't exist.
36:48 So it will just like yeah, still the we will fetch some information because there will be some table which is talking about sales, There will be some table which is talking about the revenue. There will be some table which is talking about all of this information. We'll still fetch the information and show it and which could be wrong. So how do we actually find these pitfall that if certain thing doesn't exist, how do we handle this information?
37:08 So thinking like this and you start building like around the guardrail said, okay, make it simpler for error state. Whenever these type of error occur, we show off this information to our data engineer or data analyst that okay, when this query occurred, there was a exceptional handling or I don't know, there are some error occur behind the scene which showed that the right information is not available because of this field doesn't exist. So next time when someone is asking there, we can actually hint them that, okay, if you're asking this.
37:38 And we try to translate in a fashion so we have a correlation between the customer query and the data tables, field name and everything. So thinking like in this fashion is very, very important.
37:49 Yeah.
37:49 And again, all of this system will grow with the time. And as you expose this, maybe the accuracy rate where we start with the 30%, 40% and with the time it will automatically improve and you will feel it because the jump between 40 to 80 will be very fast. And then it will slow down again from 80 to 90 and 90 to 91 going to take the exact same amount of time what took 40 to 80.
38:13 So as you progress with the accuracy, it will automatically move in a direction where like more effort required to improve the ecosystem.
38:20 Got it. Okay, so that's interesting. Okay, so let me recap a little bit. It sounds like I like how you kind of infuse that with a lot of user experience sort of design. So that, so that's not just like a black and white like hey, did we get the answer right or wrong? But you have the like, hey, initially, you have a query initially in the chat interface, you are giving all these cues about hey, this is a place to ask questions about data.
38:49 And here are successful questions that we've been satisfactory in a satisfactory manner we've been able to provide answers to. And so that kind of gives a cue to the user that this is type of the question. This are the, these are the types of questions that we are pretty confident about. Like to nudge them to like, hey, don't get too crazy in your, in your questions. And then once they actually ask something, parse that somehow translate it to queries.
39:16 And so we have to figure that part out and then pass it off to the warehouse and we have to figure that part out. And then at the end, kind of like your evals there is like, hey, give it a gradient of the extreme cases and like somewhere in the middle and kind of give custom logics to what happens when it encounters those. And then over time as more volumes, as more examples are given to the bot, then we have in place a better and better system for just over the longer run.
39:48 Is that roughly kind of what your suggestions was?
39:52 Yep, absolutely. I think that that's like absolutely what I was saying. And I think like one more point if I add this, that we can also start with the data engineer reuser only so they understand the backend in a fashion that what are the different tables exist, what are the field name exist? So also like using these type of like a better version with the data engineers, data analysts, data scientists, to start exploring like add one more layer of perfectness that okay, whenever we are getting response back, we have a query and the query was right or wrong and how the system was acted.
40:28 What are the different things they called like, what are the different responses we got back? What was wrong? Was it right? So all of these, like starting with this point makes a lot more sense than directly giving to the end user because end user, like again end user has a like very different expectation and they always think that every product is a very ideal product.
40:45 Right.
40:46 It just like start asking things which is unrealistically you will never ask in the real world scenario.
40:52 So yeah, I feel like the first couple iterations of this project now that I'm thinking about is like the goal is not to solve the this like pain point of getting these stakeholders or business users or whatever, or managers, engineers, whatever the information they need for the job. The goal I think is actually for like start collecting data on what kind of questions we can answer and what kind of questions we can't answer.
41:23 It's almost like and like, like you said do it internally first on the data team and then maybe go out and be like, hey, like this thing kind of sucks. But like just start asking questions because like we need to like figure out like where it fails and we need to like kind of set it up in a way. Like when you, I think when you have the goal of like our goal is like evaluate and see what it can and can't do.
41:47 Then you start building in the mindset of like from the very start, like what data you need to track to do that rather than with bold building and like, hey, we're just trying to get something that answers someone's question for their for like the actual like impactful use case. You might kind of skip over all this data you need to be tracking to make it better.
42:07 Agreed, Agreed. And I think as we are discussing this, I got one more insight on this is like for example, if as a stakeholder I say that I'm 95% confident that this is the quarter to insight for the stakeholders, they probably will not understand this. The executive will not understand. Okay, what does 95% means? That it's either a right insight or say wrong insight.
42:32 Right.
42:32 So if you say this to a data engineer or data analyst, they will understand okay. That most of the facts are correct, but I may need to just see if something is missing. Right. So these type of things like maybe really need a data analyst and data engineer to verify once that okay this confidence scores or cues what does this means? And then you also need to talk this that okay, I'm 99% sure that the response is correct.
42:59 Right. So adding these things will also help to, to know what's happening with the ecosystem.
43:04 Yeah, like that's a new job on the data team. It's like not even some of the work I do around AI evaluation is like evaluating AI in our product. There's probably like some role that's going to get created that's like you are the AI evaluation person for your internal data tool. Or it's even like you know how there's like like Salesforce engineers like they like just work on like the Salesforce data for within a. Internally within a company.
43:30 Like I could see it being something like that with like Snowflake or Databricks or whatever has this like AI feature that has to, that can be hooked up but like you have to have all of your underlying data and stuff and implement it in a certain way for the AI to actually like work well with it. That's probably an entire job in itself.
43:49 Exactly, exactly. Because how do you like just think about this, how dangerous it can be. So you go, you go with the data set which is comparing quarter two and quarter three sales data. And I don't know because of hallucination if the number came opposite, you go and present the same portion to the executive.
44:09 Guaranteed. I guarantee there are going to be some stuff seems shit in the future where it's like someone did their earnings report publicly. They did, the AI got it all fucked up, their stock went crazy and then they have to come out and just like oh man, like false report, CFO's fired. Like there's going to be shit like
44:33 that and they will not say anything to you. Maybe they will say it to you after the meeting, but your manager will get going to ping that. Okay. Why didn't you verify once before presenting it?
44:44 Oh yeah.
44:45 So like I'm saying like really hallucination is real. Like everyone knows this. For example, like I am. It's so real that like sometime when I'm using ChatGPT Copilot and I ask, okay, what's the stock price of Adobe or ServiceNow? They come back with the wrong response and then I tell them, okay, you are wrong. So they go back and then find the right information and still sometimes it's wrong. So that's like I'm not asking very difficult question.
45:12 It's a very simple question. But still there is a, like a very high probability to get the right. Right information versus wrong information. So yeah, like we still need someone to verify that. Okay. When we are presenting this numbers, when we are presenting this or we are finding like even not presenting, even if we are using these numbers to take decision for the next iteration or next product decision, all of those are true. It's not wrong.
45:38 So things like this. Right. So definitely it's very, very important someone to actually verify all of this.
45:42 Yeah, I feel like that's just like.
45:44 Go ahead, go ahead, go ahead. Well, no, go ahead. I was going to follow up with a pretty big bucket of questions there.
45:51 I was just going to make a point. One more other thing that now I'm thinking of is for our data agent thing that we're making that the three of us have talked about here. Yeah, it's like there's probably some balance between like, where is the area where you're getting a bunch of questions that are just like sucking up people's time, but it's also very low risk if you get them wrong. Like it's probably not that. Probably the last thing you're going to do is what we've been talking about, which is the executive like stuff to the execs.
46:21 It's probably like whatever the like lowest risk stuff. But there's like just like it's sucking up time. Try and nail it there first. Then yeah, when you're ready to do some like stuff that could. Someone important, really important is going to see your. That's going to be like in a. Yeah. Financial report externally for like a public company. You got to have it locked in before you get to that level.
46:46 Agreed, agreed. Like, and I think exactly right. So that's why I was Saying that giving them most frequently asked things.
46:53 Yeah.
46:53 Like a easier way to just like say oh I love that. Oh yeah, this, this is like my question. Yeah. Just like let's click on it and just use it. So instead of like thinking about like a very difficult. Because if you give them open ground, the probability of getting different style of question is very high.
47:08 Yeah.
47:08 Yeah, that's so interesting. Okay, so here's, here's the, here's the question I have for you Lohit, in the spirit of this chatbot that we were, that we've been talking about, call it the all knowing data agent, like what Sean just said. The spirit of it is such that for example like customers or stakeholders instead of going to a data analyst, instead of going to a data scientist, they have access to it and they can ping it whatever they want.
47:35 The one thing that worries me a bunch and we kind of talked about a lot about the mitigations here is you know, like hey, it gives up the, gives the wrong answers and things like that. There's the other side of it which is if the agent or the chatbot gives wrong answers once or twice or three times, most likely the stakeholder is going to kind of certify it to be ineffective. Meaning like hey, I'm never going to use it again and like it's game over for the spirit of, you know, hey, by like, like give us a little bit more time so that we can improve it and it will become really useful in the future.
48:14 How do you kind of think about solving for that aspect as well knowing that hey, there's risk if there's hallucinations, but there's also you only get one shot sort of first impression with any technology out there.
48:28 I think that's, that's a very, very valid question and very valid concern. So my opinion is exactly, that's what I say. That releasing beta version internally and trying to break the system is very, very important. Before sharing this to executive really deep understanding that where the system failed is very, very important because just think about this is not like again you can automate some of these things but really understanding where you got the wrong responses is very important.
48:57 There are a lot of factors you can get the wrong responses. Data was missing. The name of nomenclature of the data field was different. Defined definition of metadata was different. Fetch the right information because. But some table was unresponsive so we just only got one part of the information. And like just think about or again like hallucination or different modality. Right. Some was text, some was numbers. Couldn't find out right relevance between numbers and text. So so many ways this can actually break the system.
49:29 The ideal scenario is that like I have some like again this is. I'm going to share some strategies but these strategies like depend upon how do we improve hallucination between ecosystem. But this totally depend upon the domain depend upon the problem statement, depend upon how you're implementing and how the ecosystem works. So it's not like a one arrow which can solve everything. There are a bunch of things you need to do and that's why the data engineer and data analyst required to know where we are failing.
49:55 Because like before even going to the executive we need to just do like a smoke test to understand. Okay like if this is like a even version one like can we just answer top 15 most asked question in all different fashions Executive can change their words till we should be able to find out the responses. Very very important Second thing is that once you have the top number of queries which you are responding properly. What was the different outliers which exist when you are responding that or how do we build more trust between the responses and then a person who are using those responses.
50:30 So for example if as I encountered that when I'm asking the stock information with the copilot now I always have this notion in my head that okay the response what you're sending is not right. Can you double check the response and double check the effect. So how do we build this confidence core ecosystem that okay how confident we are that we are responding and it which is a very difficult problem to solve right now. Yeah like very very difficult problem for LLMs finding like this multi layer strategy which has model level improvement, system design improvement and user experience improvement.
51:04 All of those all together require some work before you share this with anyone. Yeah and if you want like I can share some like standard things like different people do. Like for example using RAG to just confirm the definitions and like just going on the web and finding more resources fine tuning prompt engineering like these type of things or query transparency like I don't know like how useful it is for executive. But at least like when you show the query it build a little bit more confidence in the terms of like user experience queue.
51:38 This was a query and this was that's why the response was generated. Or I don't know like. Or like putting human in loop. That's very very important. Very very important. And I don't know like if we some way we can also generate live data fact check that would be so awesome to see that you have a response and then once you get the response, there is a different system who just verify that if it's accurate. So I don't know, something like this could be very, very helpful before we go in real world and start sharing this with the leadership or executives or even like exposing to the, to the real customers.
52:15 Yeah, I think like listening to you and kind of thinking about it too while you're, while you're talking to. It's like there's, there's what you're saying where it's like in terms of instilling confidence in what you're working on. And so people don't just get scared away when it doesn't work their first chance. It's like those tactics around like transparency, just like you're saying rag and showing the query. There's having a human in the loop as like another risk mitigation.
52:38 I think those things are like around like how can we do as much as possible to like show that this is working or why it's working the way they're working. I think the other part of it is just like a kind of mindset and investment decisions like from your company, from your team, where it's like, hey, if you're going to go in and you're going to try and build this thing, like it's not going to be perfect the first time like we were talking about where like new rules will probably be created on this stuff.
53:05 And it is like. So like think about it before you go into it. If you're going to be willing to like make the investment on implement on iterating on it and maintaining it, because it's probably going to be like a forever thing that but, but it could save massive on resources to, to invest in that. So it's not like a little hacky short turnaround thing. And then I think the other thing that I think is just like a mindset for using AI in general.
53:31 Whether you're using like Chachi PT for just like, I don't know, planning your vacation or you're making a product or like using it to code or whatever is just be in the mindset constantly. That today is not the limits of what it can do. It's like, I think like Kevin Wheel was on Lenny's podcast and he said something like if you're doing something right now and it's like 60% there, it's gonna work in three months.
54:01 Like just chill, like it's gonna be good. Like my wife like a year ago or two years ago used ChatGPT. She's in digital marketing and she was trying to use it to automate some of her or help with some of her reporting and it sucked. Like it was, all the calculations were wrong. It was like back when it hadn't really liked, used Python to do math yet. And then I was like, dude, you need to use AI or you are not going to have a job in the future.
54:27 And we need dual income. And so she tried it out like a couple months ago again and she was like, whoa, it's way better now. Like it works for all these use cases. I'm like, yeah, like it progresses like crazy. Don't get like caught up just because it's like, get some stuff wrong right now. Like, even if it gets some stuff wrong, like it's still useful right now. But this is the worst it's ever going to be.
54:50 It's so much better than it was a couple of years ago. So I don't know, just going to that mindset to the, to the users also it's like, hey, if you're going to use this and you want to use AI to like automate some of your like data requests, just go into it with like, it's not going to work totally. In order to get it to work, we need people to use it and track the data.
55:06 Yep, yep. And I think like, again, yeah, this is kind of like a transition period. Like learning about these technologies is important, very, very important because it saves time. That's a very, very important part. Like really, if you think about this, I know it's like sometimes it could be daunting, sometimes it could be, oh, just think about this, like for everyone, don't need to learn everything behind the scene. That's why the engineering and the product and the companies working so hard to actually give you a user experience which is simple enough to just consume it.
55:40 So like all the companies not going to try to solve what you ask. Right? Right now that how do we build a system which can use the tables and generate a conversation? AI format insights. There will be some companies who will be working on this, focus on this and other companies probably will be using this. So all the problems, like, everyone's not going to solve all the problems. Different company will focus, have different agendas. Adobe has like a design perspective.
56:07 Right. So they will think about that. How do we use AI in design and creativity? ServiceNow is like a, like, how do we use AI to improvise the overall workflows, what different companies are using and this is the way it works. Right. So like not every company will be focusing on all the problems to be solved. There will be some problem will be solved in a different companies. And again, like, this is daunting, but this is like very, very useful if you see the positive side of it.
56:31 Oh yeah, yeah, no, absolutely, absolutely. I mean all the data warehouses, like snowflake databricks, they're all trying to come up with like text to text to answers sort of paradigm. I think they're definitely at the edge of capability right now for the LLMs. They don't hallucinate a bunch and everything that we just talked about. But yeah, I think the good reminder is it's certainly coming. If you want to build it yourself, it's absolutely possible as well.
57:00 If you want to wait for others to build it such that then you can just leverage instead of investing resources to it, that's an option as well. My takeaway is we're still need to be a little bit patient to, to get at something that's, that's, that's, that's, that's really, really, really helpful. And then certainly it's an investment. It's not a one off like, hey, let's click this AI mode and click Enable and then you're good to go.
57:24 So that's, that's awesome. Cool. Hey Lohit, thank you so much for your time. We're coming up on an hour of recording that went by really fast. So yeah, I think our audience really can benefit a lot from just hearing the product leaders who are actually deploying real agents in the world and how that would affect even the people who are doing a lot of the data work from the audience from this podcast. Cool. Sean, do you want to close it out?
57:52 Yeah. Thanks everyone for. Yeah, as Hai said, thanks for joining us. I definitely learned a lot. It was pretty fun walking through just like potentially how we could build out some sort of agent to do some of these data analyst tasks. I feel like there's a lot of people who are kind of like talking about this and talking about the possibilities, but I actually don't see that many people who are actually doing it yet. So it was really cool talking to you since you have done a lot of this stuff in the past and getting kind of your input on the pitfalls and risks and mitigations and also the solutions to those.
58:28 Yeah. And to our listeners, thank you so much for following along. If you enjoyed this episode, please, like, please comment and subscribe. It really helps get our video in front of more users like you that enjoy this type of content. Lohi, my last question for you is like anywhere that our listeners can follow along or if you have like I don't know if you, if you're on like Twitter x like or LinkedIn or if people want to follow along with like kind of what you're working on.
58:57 Yeah, so okay, so first thing, thanks for having me. It's always great to share with the community. If you guys has any, if you guys have any other question you would like to know more feel free to reach out to the the folks and it's a great, great podcast to actually talk about the real world problems and how the different communities using it. So please like and subscribe the folks and always good to share the words with audience so they can also like progress in their own career and build their own path.
59:28 You can follow me on the LinkedIn. I'm on LinkedIn so my still name is Loitak Shiyogi. You can find it. I think it should be very easy to find me but. Yeah, but anything, anything else. It was amazing talking to you guys. Great, great questions. They're amazing insights because it was fun talking to you.
59:48 Thanks so much. Bye everybody.
Summary
- Chatbots and conversational AI are similar but differ in scope; conversational AI encompasses various modalities beyond just chat.
- Early chatbots were limited in understanding context and often produced incorrect responses due to rigid programming.
- The transition to large language models has improved chatbots' ability to maintain context and generate relevant responses.
- Implementing a chatbot requires careful planning, including defining user interfaces and exposing relevant data to guide user queries.
- Human oversight is crucial to ensure accuracy and to handle ambiguities in data interpretation.
- Building trust in AI systems involves transparency about how responses are generated and maintaining a feedback loop for continuous improvement.
- Organizations should start with internal testing to identify potential failures before deploying chatbots to external users.
- The future of AI in data analysis is promising, but it requires ongoing investment and adaptation to improve accuracy and user confidence.
Questions Answered
How do I go about investing in a chatbot for my overwhelmed team?
To successfully implement a chatbot, it's crucial to have your underlying data organized and accessible. The AI needs to work with well-structured data to function effectively.
How does AI understand context in user queries?
AI systems can interpret context by recognizing relevant information in user queries, allowing them to ask clarifying questions without needing all details upfront.
What are the limitations of AI models in handling inconsistencies?
AI models can identify inconsistencies but cannot resolve them without human intervention. They require guidance on how to handle missing or ambiguous data.
How do we assess the accuracy of AI-generated information?
AI responses can be categorized into four buckets: absolutely correct, absolutely wrong, partially correct, and partially wrong. This helps in evaluating the reliability of the information provided.
How can we ensure a positive first impression of AI technology?
To avoid negative perceptions, it's important to thoroughly test AI systems internally before public release. Understanding where the system fails is key to improving its effectiveness.