Section Insights
Introduction to AI Motion Graphics
Are we finally able to create motion graphics with any AI model?
The speaker discusses the current state of AI in motion graphics, indicating that we are about 80-90% there. They plan to explore various systems, case studies, and comparisons between different models.
- AI motion graphics are approaching maturity.
- The discussion will include system comparisons and case studies.
- The speaker will evaluate different approaches to motion graphics.
Generating Motion Graphics with AI
How can AI generate consistent motion graphics?
The speaker explains their method of generating clips by locking a style and replacing action elements, ensuring consistency across clips. They emphasize the importance of style sheets and the ability to change specific elements without altering the overall style.
- Locking a style while changing actions leads to consistent output.
- Style sheets are crucial for maintaining visual coherence.
- Different models may have varying capabilities in generating motion graphics.
Challenges in AI Motion Graphics
What are the limitations of AI in creating detailed motion graphics?
The speaker notes that simpler movements are easier for AI to handle, while complex movements can lead to challenges. They highlight the trade-off between speed and accuracy in motion graphics generation.
- AI struggles with fine details and complex movements.
- Simpler movements yield better results in AI-generated graphics.
- There is a balance to strike between speed and detail in production.
Comparing AI Models for Motion Graphics
Why is Omni being compared to other models like Remotion?
The speaker compares Omni to other models, noting that while Omni performs well, Remotion has advantages in accuracy and specific data generation. They discuss the strengths and weaknesses of each platform.
- Omni excels in certain areas but Remotion remains superior in accuracy.
- Different models serve different needs in motion graphics.
- Remotion allows for more precise data manipulation and cost-effective production.
Advantages of Remotion Over Omni
What makes Remotion a better choice in certain scenarios?
The speaker highlights Remotion's strengths in generating accurate data and its ability to create assets programmatically, making it a better choice for specific use cases despite Omni's advancements.
- Remotion offers greater accuracy and specificity in data generation.
- It is more cost-effective for large-scale video production.
- Programmatic capabilities allow for efficient asset creation.
Transcript
0:00 Are we finally able to create motion graphics with any AI model? In today's video, we will talk about AI motion graphics. But before we talk, let's just watch this video together here. >> Right now, a handful of companies are holding up the entire US economy. Nvidia invests in OpenAI. OpenAI pays Nvidia for chips. Microsoft funds both sides. The same dollars run in a circle, and every lap, the valuations climb. AI data centers are now driving more growth than American consumers. But MIT just measured the gap everyone's been ignoring. AI can technically do 1.2 trillion dollars of human work.
0:35 Companies are actually paying for a tiny fraction of that. And the men at the top, they just keep pumping. Sam Altman wants trillions. Elon Musk needs xAI to win. The White House needs the boom to last. The people inflating the bubble are the same people in charge of policing it. So, when the gap between promise and profit finally closes, it won't be a correction. It'll be a chain reaction. Chips, data centers, stocks, your pension, all wired to one bet. The question was never if it pops, it's who gets left holding it.
1:15 >> So, the big question is, are we finally there? And the honest answer is, I would say 80 to 90%. In today's video, we will talk about where we stand at, like the story, cuz I already had a video 8 months ago. Then we talk about the system, and we talk about like any style test that I did, some case studies that I did, and then also some examples from other models, and some comparisons, basically. The last question is also Omni versus Remotion. So, text to video or image to video versus code. And which approach is the best for motion graphics? Okay, some background story.
1:43 Around 8 months ago, I have created already this video, if you're interested. This was basically before all the hype around Remotion. And in this video, I explained why you should not use text to video or image to video, because back then, it was really not that great. There was a lot of hype already regarding AI motion graphics, but if you look at the details, we were not there yet. Like models such as VEO 3 or Kling 2.6 were just not good enough to render basically motion graphics in a way that was consistent enough or any text at all. Since the beginning of this year, Remotion got a lot of hype, and I think Remotion is still the way to go if you want to generate bar charts, line charts, anything with a lot of detail, and not too crazy animation. And while people talk about Omniflash in video to video, because this is of course delivering some great results, I think Omniflash is actually also pretty good as you could saw in terms of motion graphics. Now, before we go to the details, two things up front: why I think Omni is so great. As of now, the quality is not that great in terms of resolution, but it has a very great understanding of reference images, and also it was very good at creating typography. And I was testing this across different styles, like not just the box style, but also blueprints or any glass animations. And I was testing this not just for this project, but it was literally something that I was seeing across all those tests. But, let me clear, we are not 100% there yet. If you look here at the clip closely, you can see a lot of mistakes. And yes, on my example, I was cutting out those mistakes that were obviously there. Of course, if you're pixel peeping, you will see a lot of mistakes here and there. And I will talk about the details as well. But, I would say in general, we are really 80 to 90% at a stage where I'm saying like, "This is usable." And if we hide this correct, and if we reach right, we can actually generate some quite good-looking results. Okay, now let's talk about the system in which I was able to generate those clips very quick, actually. Because regardless of the style I've used, the system was actually all the time the same. So, I generated one style, and then just replaced the action per clip, basically.
3:35 So, you are probably familiar of those character sheets, style sheets, and there are so many sheets, and I was trying this actually now with motion graphics as well. I created one style sheet, and then just replaced the action blocks here per clip. So, again, I had a style lock, replaced the action, so that I was staying as consistent as possible in my clips. This is basically how I structure the prompt. I was generating again like one style reference and this part here and this part below here, they stayed basically locked in the whole time unless I wanted to change the background. Now, this is important because in some shots I didn't wanted to change the background. On this vox style, I really wanted to have also some variations there as well. So, I did change actually this upper part as well.
4:17 But after trying a lot of different generations, I can really confidence say that in most cases, all you need to do now is to change your shots here. And once you lock the style, that's actually already good to go. But not with every model. We will do again a test later on where you can see there is actually like a huge gap between the understanding of each models here. Now, here some examples. In some cases, I just used a reference sheet, generated an image, and then we created the video clip. And in some other cases, I just had the reference image and generated the video clip straight from the reference image like really like reference to video. And in most cases, this was working totally fine. But in specific, if we wanted to use logos and faces, I think this way is the way to go. Here are some style tests I did so that you see basically that no matter what style you are using, it's the same workflow. So, I've tested this basically with multiple styles as you can see here. We have a blueprint style, a clay style, a vox style, and you can find those examples and like references maybe on Pinterest and other like great platforms out there. Here are some prompt examples you can take a screenshot and ask basically AI to enhance it with you. I'm also on a motion skill right now which I have used for this video and I'm not sure if I'm ready to share it already when this video will be out there, but I will share it definitely anytime soon. So, here same principles as you can see again take a screenshot and then try to ask AI to regenerate something like this for the specific style you are looking for. Here of course like also this vox style and as you can see we have here really the elements that are important.
5:44 We have the colors, so what colors do we want here? We have somehow the font I haven't specified the font but you can see it's like caps lock and then you have the elements we want to use here. Again, like you can see it all in the prompt written out here, okay? And that was really fine enough for Omni to understand already the entirety of the style. Okay, here a rule of thumb that you can basically apply if you are generating motion graphic clips. If you have big fonts then AI is struggling less than if you are looking here at those micro tiny fonts here. So on this side you can see while it looks quite good already but if you are looking very close of course the text here inside some of it is not accurate and even here but again this is what I mean with pixel peeping because it's actually already quite good but you can see here on the left side obviously that's great looking. The font is great, the characters are great. Of course we can now argue about the background but even then it's fantastic actually and I think overall it even understand logos and not only that but if you have like simple movements and like character faces it do understand it but we have to understand actually that Google Omni is quite sensitive when it comes to known faces.
6:47 Now there are some workarounds but we will talk about them in a bit. In general again try to avoid fine typography and micro details like on this right side here and in terms of the movement I think you can guess it as well. The less fancy the movement the less the AI is also struggling. So keep in mind you want maybe to create rather simple movements if you want to have longer clips or if you want to have dynamic clips where it's okay to just have a specific bits and pieces then you can be going a little bit wilder and do some nice dynamic shots. Here are some examples of really fine grain. They might be too small here on the screen but here you can see that we have here one kind of similar background. The style obviously is kind of similar so here we have this style which is coming in my next video but here you can really see that Omni is doing really a fantastic job and yes we could do a lot of it also inside of Remotion but it would take us so much more time and that's something you have to wait. What's more important? Speed versus accuracy versus also dynamic because doing a shot with such a high dynamic and a lot of movement a lot of details is also not that easy in re-motion. All right, let's talk about the case study here. So, the Vox Explainer. There is one thing that I noticed right away. Those tests wouldn't be possible with today's sponsor Artlist. They sponsored me on this video here to basically run a lot of different tests. And as we can see here, I was really generating like a lot of different tests, failed, tried, and perfecting basically the prompt structure here for different styles with different providers. And I was really trying all kind of different stuff here.
8:11 One thing to notice here, inside of Artlist, they have Gemini Omni Flash. But besides that, they have actually a lot of models that you can use now technically for free. So, they have this unlimited option here that allows you to generate images such as on Google Nano Banana or GPT-2 without any limitations. And that's fantastic if you want to try around like styles like how are those reference sheets looking like and so on. All my reference sheets were generated with their unlimited option here. Now, going back to the case study, there was one thing or one incident that Google Gemini constantly did not generate images at all. Every time I wanted to bring in some known people, that's where Gemini blocked me basically. So, what I did, and this is why you see here like those images with those black bars, that was actually by accident. Overall, I think I like it. I would prefer it even now with those cuz I think it makes it even look nicer. But it again, it just happened so that Gemini didn't let me to render those video clips. And then I asked GPT to create those images for me first and then put like a black bar on top of it. So, it was like a text to image first and then an image to image just to apply those black bars basically here. So, the entire clip was basically a mix between reference image to video and reference image to image, reference image here to the image itself. So, this was like the image I got first and then I regenerated it with those black bars because that was like the safest way to do it. But regardless, I was always using and applying this reference sheet to generate basically everything from the film. This is actually the most important sheet. And when it come to logo, I did the same thing. So, again, I took this reference image and then I generated an image first. So, here is the prompt for the image and then I generated out of this image here the video prompt. Now, this is also very important. Even though you are attaching this image here, you have to specify in a style section here that you want to have the style. I tried to also leave the style block here and it was actually not working because it transformed style mid video. So, again, be very specific here to lock the style also in your video prompt so that the style is consistent here as well. Okay, I'm just sharing the video now with every single prompt from every single clip so you can see it. You can always stop it. You can always take a screenshot and reverse engineer it yourself. So, here we go.
10:28 >> Right now, a handful of companies are holding up the entire US economy. Nvidia invests in OpenAI. OpenAI pays Nvidia for chips. Microsoft funds both sides. The same dollars run in a circle and every lap the valuations climb. AI data centers are now driving more growth than American consumers. But, MIT just measured the gap everyone's been ignoring. AI can technically do 1.2 trillion dollars of human work. Companies are actually paying for a tiny fraction of that. And the men at the top, they just keep pumping. Sam Altman wants trillions. Elon Musk needs xAI to win.
11:06 The White House needs the boom to last. The people inflating the bubble are the same people in charge of policing it. So, when the gap between promise and profit finally closes, it won't be a correction. It'll be a chain reaction. Chips, data centers, stocks, your pension, all wired to one bet. The question was never if it pops, it's who gets left holding it. >> Okay, but why Omni now? Let's do this four models comparison, and I know it's not a fair comparison, but let's still talk about it here. Okay, what you can see right away here on the left side here, we have Kling 3.0, and yes, the first frame is our reference sheet, and as I've mentioned, this is not a fair comparison. Why? Because Kling does not has a reference mode, at least not inside of Artlist, but they have only a start frame mode here. So, that means for Happy Horse, I was able to generate just reference SeaArt 2 as well, and of course, obviously, Omni Flash. Now, in terms of the understanding, we can already see Flash is really the best understanding. It's exactly like how I want it to be in terms of the fluid animations, and you can see Kling is also struggling here in compared to the others. Happy Horse also not that great.
12:19 I would say SeaArt 2 looks good, but if you are going closer, you can see the numbers are really not great at all. The people are not great. You can really see that Omni is by far the best model here. Here on the other style, we can see the same problem here, of course, with Kling, and then we can see SeaArt on the surface looks okay, but when you look actually closer here, it's transforming the camera into a camera. And while on Omni Flash, it's also not 100% there, I would still say, in terms of understanding how the smartphone is rendered, it's actually the best here.
12:48 And also, in terms of the typography, we can see that actually it's not falling apart, but everything is readable. And Happy Horse, it's not great at all. It's, yeah, very bad in compared, and the Kling 3.0 is here not usable in any way. Now, this animation here was less of typography. I just wanted to see like how the physics are, how do the animation behave, and here I still prefer Omni Flash, regardless if it's like a simple animation or not, but SeaArt actually is also doing good here.
13:15 So, overall, I would say really Omni Flash is the best in almost everything here. Of course, when it comes to generating like cinematic images, I would still say SeaArt 2 is the best model, but when it comes to like video to video, physics, and also like motion graphics and understanding text, I think Omni Flash is actually beating Cidence 2. And not only that, but it's also by far the cheapest model from those like top-tier models here. I would say where it still could be better are the details and sometimes the motion. But again, I wouldn't give any of those models a full point here, but I think in terms of understanding, it's actually already doing a really good job. It's just the execution could be a tiny bit better.
13:52 Now, let's talk about Omni versus Remotion, because I think that's also a topic that might be interesting. I think you have a great feel now what Omni can do, but of course, like everyone is still talking about Remotion, and I still believe Remotion is fantastic there. Now, let's go to this platform here again, and I have been talking about this platform already for a while now, but still it's like a platform where you can generate like free motion graphics completely for free. And here some examples where I believe like Remotion is actually better than Omni, and it will stay better, even though Omni can get better and better over time. But there are still a lot of use cases where I believe like Remotion still makes more sense here. First of all, accuracy. Like the quality and the accurace, being specific about some data, that's something it's like almost impossible to generate like from text to image here. Of course, you have something that looks almost perfect, but if you want like very accurate data, so for example, let's just say on February, we had a bigger spike here. And now let's increase the data here. As you can see, I can really fine-tune the data here. So, we can increase it here a bit more, a bit less, but we can be very specific. And not only that, but of course, it's also cheaper, and we can actually also do things programmatically. So, let's just say we want to render a thousand videos and want to say hello name XYZ. Now, we can just replace the name, same as we can replace the values here. And now we can generate like a thousand assets for literally nothing. So, I hope that's very clear. So, we can do stuff programmatically, but also very accurate. Obviously, there are also a lot of limitations here. So, we can't do like those Vox animations with fancy camera movement as easy then we would do it like just using text to prompt text to video basically or image to video.
15:28 And of course, we could also work with images and try to recreate vox animations here using re-motion. And for simple parallax effects and in specific now with the new re-motion effects here, you can really do all kind of controls there as well and change images, backgrounds, and so on. But I still hope you understand where the strength of Omni Flash versus re-motions are. In a nutshell, if you want to generate anything with numbers, charts, really fine typography, or like you want to scale your generations and you want to save money, I think re-motion is still the way to go. But if you want dynamics, easiness, really quick results, and you want actually simplicity, then Omni Flash is obviously the way to go here.
16:07 But I think it's not this or that. I think it's really the hybrid approach that wins at the end of the day. Because if you're able to use re-motion together with Omni Flash together with After Effects, of course, like you are really having the best of every single world here. Because at the end of the day, it's just like understanding what each tool can do and trying to get the best out of them. So if we summarize everything, then first lock your style.
16:30 Generate a good-looking reference image because we have finally a model that understands them. Try to think in big blocks because if your animations are so micro detailed, then the AI might still struggle. Sometimes you don't need to generate an image at all because the reference image is fine enough. So you can skip the step entirely now. And keep always in mind if numbers and accuracy is that what you're looking for, then we are still not there yet. But we are very close. All right, I hope this video was helpful and I'm working on this motion skill. Maybe it's already in the description. If not, it will come anytime soon. See you till the next one.
Summary
- AI motion graphics are nearing usability, with models like Omni Flash and Remotion leading the way.
- Omni Flash excels in understanding reference images and typography, while Remotion is better for detailed data visualizations.
- The speaker emphasizes the need for a hybrid approach, using both Omni Flash and Remotion alongside traditional tools like After Effects.
- Style consistency is crucial; locking in a style can improve the quality of generated clips.
- AI struggles with fine typography and detailed animations, suggesting simpler designs yield better results.
- The video includes case studies and examples demonstrating the capabilities and limitations of various AI models.
- The importance of prompt structure and reference images is highlighted for successful AI-generated motion graphics.
- The speaker notes that while AI can generate impressive results, accuracy in data representation remains a challenge.
Questions Answered
Are we finally able to create motion graphics with any AI model?
The speaker discusses the current state of AI in motion graphics, indicating that we are about 80-90% there. They plan to explore various systems, case studies, and comparisons between different models.
How can AI generate consistent motion graphics?
The speaker explains their method of generating clips by locking a style and replacing action elements, ensuring consistency across clips. They emphasize the importance of style sheets and the ability to change specific elements without altering the overall style.
What are the limitations of AI in creating detailed motion graphics?
The speaker notes that simpler movements are easier for AI to handle, while complex movements can lead to challenges. They highlight the trade-off between speed and accuracy in motion graphics generation.
Why is Omni being compared to other models like Remotion?
The speaker compares Omni to other models, noting that while Omni performs well, Remotion has advantages in accuracy and specific data generation. They discuss the strengths and weaknesses of each platform.
What makes Remotion a better choice in certain scenarios?
The speaker highlights Remotion's strengths in generating accurate data and its ability to create assets programmatically, making it a better choice for specific use cases despite Omni's advancements.