Transcript
0:00 This is the laptop that always makes me feel kind of stupid because normally since I have this YouTube channel, I have the Max version, just the M4 Max, and now I have the M5 Max, which is supposed to be the ridiculous one. And then I go back and test the Pro chip, open the same apps, do the same work, push it through the same nonsense, and I start having some very uncomfortable thoughts. So, I decided to test everything.
0:22 Gemma 4 E4B. Yeah, it's new. So what? MLX can run it. This is an 8.2 GB model. fits on all three machines. Token generation on M5 at 25 tokens per second, M5 Pro at 51, and M5 Max 87. Woo, there's a difference. But what happens when we step up to a bigger model? Quen 3.5 27B at 4bit quantization, 15 GB. The M5 Pro runs it at 17.6 tokens per second. The M5 Max at 25.1. Both usable. The M5, it crashed.
0:53 And then there's this Quen 3.5 122B, a mixture of experts model at 65 GB. The M5 Pro maxes out at 64 GB of RAM. And this model is 65 gigs before the OS even loads. It literally cannot exist on this machine. The M5 Max with the 128 gigs loads it fine, generates a 23 tokens per second. This is the why you buy the Max, not because it's a little faster, because entire categories of models simply don't exist in any other
Summary
- The M5 Max outperforms the M5 Pro and M4 Max in token generation speeds.
- Token generation rates: M5 Max at 87 tokens/sec, M5 Pro at 51 tokens/sec, and M5 at 25 tokens/sec for smaller models.
- The M5 Pro struggles with larger models, crashing during tests.
- The M5 Max can handle larger models (like the 65 GB Quen 3.5) that the M5 Pro cannot due to RAM limitations.
- The M5 Max's 128 GB RAM allows it to load and run more complex models effectively.
- The key advantage of the M5 Max is its ability to run entire categories of models that are incompatible with the Pro version.