Section Insights
The Human Side of 11 Labs
What is the origin story of 11 Labs?
11 Labs was founded by two childhood friends, who were inspired by their experiences growing up in Poland, particularly the limitations of audio narration in foreign films. They recognized a significant opportunity to improve audio experiences across various domains.
- The founders' friendship and shared background played a crucial role in the company's inception.
- The monotone audio experience in Polish dubbing highlighted a gap in the market for emotional and varied audio narration.
- The founders aim to create technology that allows for expressive and contextually aware audio across languages.
11 Labs' Model Suite and R&D Prioritization
What models does 11 Labs work on and how are they prioritized?
11 Labs began with a text-to-speech model that understands context and emotion, followed by speech-to-text capabilities. They have since expanded to real-time streaming models and conversational experiences, focusing on breaking down language barriers.
- The development process started with text-to-speech and evolved to include transcription and translation.
- Real-time interactive experiences were introduced as technology advanced.
- The company is focused on creating a comprehensive voice engine that enhances user interaction.
Voice Agents in Customer Interaction
How are voice agents being utilized in customer service?
Voice agents are streamlining customer interactions by allowing users to speak directly to agents instead of filling out forms. This has led to richer data collection and improved customer experiences.
- Voice agents simplify the customer inquiry process, making it quicker and more efficient.
- They provide valuable insights into customer needs and use cases.
- The technology is being applied in various sectors, including government services.
Integrating Technical Talent in Non-Technical Teams
How does 11 Labs optimize team structure for impact?
11 Labs maintains a flat organizational structure without titles, allowing for flexibility and rapid growth. They also introduced a scoring system to streamline negotiations in sales.
- A flat structure encourages collaboration and innovation across teams.
- The scoring system simplifies decision-making in sales negotiations.
- Bringing technical talent into non-technical teams enhances overall efficiency.
Challenges in Emotional Interaction and Audio Production
What challenges does 11 Labs face in emotional audio interactions?
While voice agents perform well in structured settings, they struggle with true emotional interactions. Additionally, while they can produce good music, they have yet to achieve top-chart quality.
- Emotional interaction remains a significant challenge for voice technology.
- Current audio production capabilities are improving but not yet at the highest artistic levels.
- The focus is on creating impactful products that align with customer needs and revenue generation.
Transcript
0:02 So, I love line charts and bar graphs as much as the next guy, probably more. the story of 11 Labs is also interesting from a human perspective. But, is you started a company with a childhood friend. So, maybe take us back to 2022 or earlier and just tell the the human side of the 11 Labs story to start. The I have the I have the most luck in the story of 11 Labs because well, it started in 2022 it felt feels like it started 17 years ago when I met my my co-founder Piotr.
0:33 all the names in Polish are complicated luckily for for for us, but we we met in high school, became best friends, took all the same classes together, and then through the years did everything together. So, we traveled together, studied together, worked together, and time is on our side. We are still best friends. It's working it's working out. and part of what started 11 Labs is is is inspiration from where we are both from. We are both from Poland.
0:57 so, I was of Warsaw. And there's a very peculiar thing in Poland. If you if you watch any foreign movie in Polish language, all the voices, whether that's a male voice or a female voice, get narrated with one single character. So, as you can imagine, pretty terrible experience. You have literally one voice narrating everything. it usually also on purpose is kept in monotone. So, you are meant to interpret your own emotions for that content. And while we grew up with this, this is still happening today for majority of content. And that kind of opened our eyes into one of the clear things across the domain across audio domain across the future will be this ability for everybody to speak any language with the same emotion, the same intonation.
1:41 and we started diving deeper into that problem and realized the problem of audio exists in so many other domains, too. Whether that's narrating the content around us, whether that's the books not being available in audio form, whether that's the and articles that we could read, whether that's that language barrier or in the future as we heard in the previous conversations, the future where human noise the robots are around us, the voice will be the primary interface to a lot of that technology and and something we would love to fix and solve.
2:09 Excellent. and 11 Labs builds frontier models for audio. I think there's a paradigm now where to build a frontier model, you have to start with hundreds of millions or billions of dollars and then figure out the rest later. 11 Labs did not take that path. Can you talk a little bit about your approach towards building this company, why you have this hasn't been replicated, is that even possible in 2026, etc. Yeah, that goes I think that continues that great luck and timing because we started in 2022. For those of you working in the domain at the time, that that was the year of crypto and metaverse. Nobody was still working on the AI side.
2:51 even further, people were starting to work, of course, on the text models and the visual models, but audio as a domain was still considered a big niche with so few researchers in the space working on on that work. So, for us that was a good part of picking that domain where A, we were excited about where that future is called, we felt that people around just didn't realize the value of that domain, but three, the requirements of what you needed to solve were very different. The audio models were smaller, so you don't need as much computer as you need for some of the other sister domains.
3:23 The data needs are big, but while there's a lot of audio data, we knew that the thing to actually get that audio working, you will need to figure out how to transcribe a lot of that data and annotate a lot of that data, which we knew we can do. And then ultimately it all boiled down to architectural side of can we can we solve that part in a good way. And here my co-founder is one of the smartest people I know and and a great researcher and has been able to assemble some of the best people in audio to to help us. and we took a slightly untraditional approach at the time. We started we started in London.
3:57 We had a lot of people between London and Warsaw and started a company in remote completely remote way. So we wanted to hire the best researchers wherever they were. We were going through the classic GitHub scraping and and and trying to reach people based on their work instead of based on their presence and based on that work we would reach out to those people. We would always share our examples and try to get them to join the team and that's how we assemble the first the first set of of of people who we think are some of the best researchers in that audio domain and through the years they still help us crank a lot of those models in into production.
4:33 then we launched the product. I think the slightly different approach we took was monetizing very quickly. So trying to get some of the revenue stream back so we can fund a lot of the work and the models. We try to stay stay healthy on the margins so we can continue investing with the assumption that it's better for us to figure out that stream and be be able to be independent in that development. but then as the ambitions grew we knew that we needed to train models so we of course brought a lot of money externally as well. And I think like projecting to today one thing that's clear for us is there's still so many of those niches that people don't tackle that that you can start with and then step-by-step start opening them them up.
5:17 I think a lot of customers see 11 Labs through their narrow needs, right? maybe take a zoomed out view. Like what is the suite of models that 11 Labs works on? How do you prioritize them? How do you organize R&D? Etc. Yeah, so we started we started with the first text-to-speech model. So the model that could finally understand the context of what's being written and based on that context context understanding get the right emotion the right intonation from text.
5:45 So if it's a happy sentence you get that happiness out. If it's a it's a dialogue, it can pronounce the dialogue out. And then continuously started adding that. So, it started with the problem of of breaking down language barriers. the things you need to solve dubbing is transcription, so understanding, then the the translation, and then text-to-speech. So, we first solved text-to-speech. Then we knew we needed to add the other component, which is speech-to-text and being able to transcribe content in a in a great way.
6:14 Then how we combine those models together. So, that's kind of was the first three models in the first first couple of years. And then of course the other things started happening across the space, which is that a lot of the reasoning models started becoming quick enough and smart enough at the same time, where you could imagine those interactive experiences being possible. And that's where we started launching our more of the real-time streaming models across audio, and then combining those into conversational experiences. So, we added effectively all the stack, all the turn-taking and orchestration to create a voice engine for a voice for a voice agent. and then on the other side as we realized that the emotionality is something we can solve, we added some of the hardest modality in in in audio, which is music and being able to produce music. So, today we spun entirety of the research of audio, where it's text-to-speech, speech-to-text, combining those models together in both localization with dubbing, with orchestration with voice agent voice engine, and then and then being able to do that across music as well.
7:14 And what's the all those things and all that interesting development work was there any oh wow moment in terms of what these products are capable of that you can you can remember? Yeah, there's so many and it's kind of the bar changes for all of us. The first moment for us was well, first moment for us they always use my voice as a testing voice because it has this weird accent. And and the first time was like when when we could replicate my voice based on a good sample, that was like a first wow moment to to myself and you always go for this moment like this is not how my voice sounds like and then you listen to yourself side by side and it's like definitely how it sounds like.
7:51 unfortunately, then the the second moment was where we first got it to laugh and people were like, "Okay, this is actually the thing that that makes the the whole experience more human. The the laughter, the pauses, the ums, the ahs, the imperfections." So, we started getting those out and those were the moment for us because we made it to the top of Hacker News with the first AI that can laugh a model, which was a very proud moment for us. One of the I feel like pinnacles of the voice performance, Matthew McConaughey giving his newsletter and his iconic lines in in in Spanish and Portuguese, where for the first time his family who speaks that language could hear him speak those languages, too. but for more recent pieces, the two two ones that we are excited about bringing to production, I think the first one is finally figuring out the emotional intelligent in that interactive experience. So, in the voice agent experience where it doesn't only get the right intonation emotion, but can understand the other side. If somebody is stressed, it gets and delivers that that that soothing, reassuring emotion. If if someone is excited, maybe it matches that. If someone speaks slowly, it makes sure to slow down. And that emotional intelligence is something that we are finally seeing internally a path to solving, which will be just a a a a continuous step change to to what's possible. And then the second one, which will apply there, but also apply into general audio spaces, audio general intelligence where you can combine audio models together in one stream. So, you could theoretically have a model that narrates, then pauses and let's say starts singing with that same continuous voice. And that's something that's extremely hard to combine today and and something that would be would be possible, I think, very very soon.
9:41 And voice imagine, you know, voice agents. And it seems like everybody is at least on the customer side, everyone's buying a voice agent. and I think intuitively you think customer support, you know, the old phone tree replacement. what's actually going on in the world of voice agents? And what do you think are are the most interesting overlooked opportunities, spots where startup founders should focus? Yeah, the of course the customer support is probably the one that everybody heard and and knows about very very well. I think the second thing and the second thread we are seeing is increasing shift to revenue-generating opportunities where voice agents can act in sales, whether it's inbound or outbound sales of of sales. It doesn't replace the entire experience, but takes and amplifies part of that experience.
10:24 maybe a good example is Deliveroo, where Deliveroo will have voice agents that contact the restaurants to capture their opening times. And based on their opening times, they can update the riders and drivers. And of course, the people ordering on when to get to to that work. All the way through to the inbound sales, where increasingly people, that's a good example of Deutsche Telekom, will be contacting to inquire about the service, inquire to buy a a product. And instead of going through the drop down, instead of going through the form, you can speak with the voice agent to leave that information. we do it ourselves, too, so we have a good metrics of an understanding of what's happening there.
11:02 One, of course, so much simpler and quicker to go through instead of going through that form. But the second thing that started happening in that inbound sales flow is we we had a lot more information that people started leaving because they would speak about the use case they are coming with, but then where it's not working, where it's working, some of the other use cases that they are evaluating. Which we can combine and then just deliver such a much better experience afterwards. On the overlooked side, I think my favorite example there's the citizen support education and healthcare will completely change. On the citizen support, like all of us would would benefit from just generally better government access, whether that's understanding how to fill in the taxes that I think many of you went through earlier this this month, all the way through to just learning how what is the policy for travel abroad and and and and and how that might affect the the the space. We've recently seen that work deployed in government of Ukraine. We think is like one of the most advanced governments on that front.
12:03 We traveled to Ukraine working with their team and what they are trying to solve is they they have a a government app which every citizen can access and get information about what's happening, but given the war, given the the the frontline and lack of that access, they wanted to figure out a new channel for people to be able to call in and get that information. So, they created voice agent effectively where you can where you can call in and get the information about what's happening on the frontline. You can get education help and some of the lectures delivered to your to your kids all the way through to proactive engagement about about staying safe and staying staying out there. And maybe last example on education front, and that's probably my favorite one as I think about that changing, it's it's just how incredible would it be to have a someone that is an incredible teacher available 24/7 we can ask him questions whether it's Karpathy all the way through to Richard Feynman and and you can learn physics with them on the headphones while you are teaching that subject or learning that subject. And and that's something that we are seeing pockets of. Like a great example is MasterClass where MasterClass of course collaborates with incredible teachers to deliver static lectures, but recently they launched an interactive version of that. so, for I don't know if that will be a good reference for for this this audience, but we we recently worked with them on bringing Gordon Ramsay that can teach you cooking. So, while you're in the kitchen he can shout at you effectively to get to get better or maybe a better one there is a Chris Voss where you can of course learn negotiation but you can learn by negotiating with Chris live on the phone to to to to get better which I thought was a phenomenal subject. Having negotiated against Mati a number of times around financing rounds, I understand now. I think it helps you to say this but I think the opposite opposite is true.
13:55 ask more questions I want to save time for the audience as well. maybe one as Constantine mentioned more than 100 million of net new ARR in Q1. Obviously the business is going very well. and you're sort of pioneering the startup founder building a foundation model applications. any counterintuitive lessons about building a company in this era that for the founders in the audience they may want to take home with them. So we are just for reference we are just over 400 people over 400 million in revenue but still keep the teams extremely small so it's it's like rough arbitrary a little bit cop is is less than 10 people is for each of the the research product even the the go-to-market ops talent teams are all smaller than that size.
14:41 most of people will have 10 direct reports or so it keeps it relatively flat and allows us to move move a little bit quicker. One thing that we've done which is in this model and very surprisingly this is very similar model that we've seen actually with the government of Ukraine. Each of the teams even the teams that aren't technical teams will have engineers within them. So our people team our go-to-market team our legal team will have an engineer in that team that helps to build of course automation upskill uplevel the the rest of the people and recently that really helps because as I'm sure many of you are going through everybody will be wipe coding and coding a lot of the the help even if they are not technical so now that kind of shifted the responsibility not the not the responsibility but shifted the requirement of how good the review needs to be for all of that work. Whereas security infrastructure implications, you will want to make sure that the output is right.
15:34 And everything on the engineering side you can put that expectation on the non-engineering side the the ability to do is that is is relatively hard. So that technical research in those teams helped us a lot to to figure this this out. And and in general there's just so many incredible work you can do by having that whether that's the scraping on the hiring and recruiting front or analyzing what worked in the past to improve in the future, whether that's upscaling the legal team on how to use those tools and then figuring out ways of recently introduced a scoring system for those on the go to market on the sales side. You frequently will end up in this negotiation with your sales team of can I give indemnity provisions? What's the liability cap? Can I give this set of clauses? And then you kind of need to draw the line of how many things you give. And I ended up being in so many of those conversations that we gave already a lot or we didn't. So now we introduced a scoring system that you can give per per size of the customer you can just give a few of those points out and in.
16:33 we just made it so much easier and of course that's fully automated now with with with how we work across that team. So that was one of the unintuitive small teams bringing technical talent in the non-technical teams, keeping relatively flat. We also have no titles which allows us to to to bring people and and really optimize for impact that they are having and then you can grow as quickly as as as as you want. The tenure will not define this and and many more. So we'll see. It's four years old company. So we'll see if that's helps.
17:04 Any questions? Oh no. Okay, Sonia. Are you seeing people deploy voice agents to actually negotiate on their behalf? And then when you are you starting to see agents actually negotiate with agents? and sorry, I I do three-part questions. when that when that world happens, do you think the agents are actually talking to each other the way that humans talk to communicate and negotiate, or do you think it's beep boop beep boop? Do you think it's, you know, it's all done instantaneously?
17:34 Like how how's how's that world going to look like? So, one early inklings of that, we haven't seen any truly successful on the negotiation front. It was like more, you know, kind of order taking. What's the price? Can we capture that? And then kind of goes back to the team. So, not real negotiation. But, there's there's few startups that we see especially on any any like organizational shifts of can I organize this event, calling calling a lot of places, getting the price, and then calling again with like our budget. So, that is happening, and I think this will shift. I think it emotional intelligence will like This is the big part that will start being important in a lot of that work, where it's not only the content that matters, but how you deliver, when you pause, that work. And then maybe the extreme version of that, which agents are are not like most of the people wouldn't do it, and and they are not good at that, is today you will see a lot of interruptibility built in, where human can interrupt the agent. But, with negotiation, you also want the opposite, where agent will interrupt the human.
18:31 It's kind of the extreme version of that. On the second part, on the agent-to-agent part, Some of you might have seen this the the we did a hackathon over a year and a half ago, and that was exactly the case where agent was speaking with another agent. They detected that they are both agents, and they swapped over to the to the different language, and that's like a more of a more efficient transmitter of information than just the the classic spoken word. And I think this will happen. I 100% like the the the big question will be really voice, will it be other transmission of information?
19:08 it depends truly on what the infrastructure is built for, and I think this will define that that experience. Yes. Okay. So you're going to catch box. Hey. >> >> Curious how you're thinking about the need for voice in a future where agents do more and more of the work. So basically what are the kind of use cases maybe where human conversation I think it's more of a follow up to the last question. Like first you all of us will have so many different devices around us and step from that you will have robots around us. So of course voice will be such an important interface to to instruct and and be able to interact with those those those those devices. In many ways I feel like the you know we see a lot of developments of of intelligence but then the the real bottleneck of the future will be how we communicate with that intelligence and I have voice and visual part will will be a big unlock to be able to actually get the most of that intelligence value in those settings which which which which isn't yet possible.
20:11 But on the flip side it's a it's yet the value of the human to human interaction will only increase. So like the the whether that's the events like this one whether that's events with your favorite artist will will increase in value with that ability of having voice all all around you. But the trust will be such a big part and something we optimize for like in between the agent and human of you know in the in the future where all of you will all of us will have a voice agent for example to call and book a restaurant or give information to a health care appointment.
20:46 All of that will require such a high degree of trust that this is you and and and authenticated user. There'll be like a level of encoding and decoding for real then encoding decoding for water might opted in human and then by default everything else will be fake which is kind of the opposite of how it is today. You detect for AI but you will detect for real authenticated AI in the future and assume it's fake. if you could pass it that way.
21:17 Andre spoke earlier about jagged intelligence. Do you see similar odd places in audio where models are good and the bad that you might not expect and yeah, what are they? The there's still so much on the on the on the bad side. I think the you know, like we spoke a little bit about where we see the voice agents working. So like this combination of the models together and support settings works really well, works reliably. And early sales starts working, but like the moment you start swapping to a true emotional interaction, not yet working.
21:51 It's it doesn't get the emotion that that well. It's slightly too slow. so that is still like I think a big step change that should work. same will apply on a on a very different domain on on the music side. I think in the music side you you can get you can get good production music. You cannot get top charts music even with artist input. I think this will change over the next over the next year or two. Can I ask a follow-up on Of course. Andre's take was that the the the reason for that was that the labs were basically training for the stuff that had economic value. Where you're training your models, is that true of you? Are you basically training for the things that make the most money or is it that there are some challenges that are genuinely harder than others?
22:35 The I don't know. We we try to train the models, build a product and an ecosystem that will drive of course the biggest impact for for for for all our customers, all users, which should correlate of course with the revenue in the long term. So like that long term perspective, that's it's going to be like minimal over the next few years, so not next year. so frequently we will train the models that might not provide that value in the short term. Or even step before, we'll like I so much time labeling the data, not only the what of audio, but also how of audio. Like what emotions did I use?
23:07 what is my voice described as? What is this music described as? So we assembled a team of now thousand plus people that have been voice coaches, musicians, artists before that can help us annotate that behind the scenes. And that will not provide value in the next 6 to 12 months, but we think it will in the next 10 to 12 to 24. and then you of course need to collect that data, which frequently just isn't that accessible as well.
23:31 Hey, last one and then we'll go to lunch. Hey. Can you hear me? Thanks. Big fan of yours and 11 Labs. Thank you. What do you think from from the model air perspective, what do you think are the moats here with with audio models? The labs are going there, not going there. What are the kind of, you know, in this sausage making of making a real good frontier audio model? What what are the the main defensible parts there?
24:03 At the So of course we do a a variety of models and recently had a pleasure of meeting Jensen and he was commenting on a few of those models and he said that our speech to text our speech to text models are technology and text to speech is artistry and we are all artists. So he gained a client for life. But of course we do believe there's a little bit of that to to really fix text to speech and fix that emotionality. You you you you will need to be really focused on that space. You really need to get in front of users, collect the data, collect the preferences, use that to fine tune the models and then there is a domain specificity in how you actually bring those models to production. In health care very different than in financial services, very different than in education or experiences. So that's on the model layer. I think there will be continuous advantage that if you actually care about the quality, that like actually spending the time on the model work will will will help you keep that advantage. But to your point, the models and like a lot of use cases will use a model as just a small part of their stack. And that's where we spend a lot of time and like beyond going beyond the research on on the product side of how you understand a user's problem, the workflow that they need.
25:12 invoice agents is combining the audio models with knowledge and bringing that inside of the the the system, how you bring it outside with telephony systems so you can interact across channels, how you evaluate, test, and monitor. and then as you create, whether that's in the agent space, whether that's in the creative space, that's some understanding, you build the ecosystem. And that's what we hope to build across 11 labs, a place where whether that's distribution and brand that people can trust, the platform where you have pre-existing set of work that you can start off, whether it's a template for creating an agent, template for creating a workflow in a creative space, or whether that's a voice. And we had a pleasure now of of having over 20,000 voices that people created, contributed, that you can you can use across language styles and voices. And I think that will be an increasingly important layer of how you are able to cater to that diversity, make it easy for people to start, and really understand that that workflow.
26:06 All right, I'm going to hand it back to Konstantin Mali. Thank you. Andrew, thanks for being partner. Yes. >> >> Amazing. Thank you, guys.
Summary
- 11 Labs was founded by two childhood friends from Poland, inspired by the limitations of audio experiences in their home country.
- The company focuses on creating advanced audio models, including text-to-speech and speech-to-text, with an emphasis on emotional intelligence and contextual understanding.
- They adopted a remote-first approach, hiring top researchers globally and maintaining small, agile teams to foster innovation.
- 11 Labs monetized quickly to fund development and maintain independence, while also exploring various niches in the audio domain.
- Key products include voice agents for customer support and sales, with potential applications in education and government services.
- The company emphasizes the importance of emotional intelligence in voice interactions and aims to create seamless audio experiences across different contexts.
- Future developments may include agent-to-agent negotiations and improved human-agent interactions, focusing on trust and authenticity.
- The founders believe that the future of audio technology will involve a blend of artistry and technology, with a strong focus on user experience and workflow integration.
Questions Answered
What is the origin story of 11 Labs?
11 Labs was founded by two childhood friends, who were inspired by their experiences growing up in Poland, particularly the limitations of audio narration in foreign films. They recognized a significant opportunity to improve audio experiences across various domains.
What models does 11 Labs work on and how are they prioritized?
11 Labs began with a text-to-speech model that understands context and emotion, followed by speech-to-text capabilities. They have since expanded to real-time streaming models and conversational experiences, focusing on breaking down language barriers.
How are voice agents being utilized in customer service?
Voice agents are streamlining customer interactions by allowing users to speak directly to agents instead of filling out forms. This has led to richer data collection and improved customer experiences.
How does 11 Labs optimize team structure for impact?
11 Labs maintains a flat organizational structure without titles, allowing for flexibility and rapid growth. They also introduced a scoring system to streamline negotiations in sales.
What challenges does 11 Labs face in emotional audio interactions?
While voice agents perform well in structured settings, they struggle with true emotional interactions. Additionally, while they can produce good music, they have yet to achieve top-chart quality.