Single post
jump to repliesThe #Apple Mac Studio Ultra at 96GB is not available in a reasonable timeframe. If ordered today in June…will not get until October. The demand is still unreal. Very irritating.
The DGX Spark is tempting… this memory crunch is peak annoying.
#selfhosting
5 replies
back to top@bryan I simply wouldn't buy anything until the memory crunch is over. Renting is a much better deal. I bought a Strix Halo with 128GB right before everything exploded (I paid $2100, I think, which is more than they first launched at, but not much more), but even that isn't that great of a deal. $4k for an ASUS GX10 or $4700 for a DGX Spark buys a lot of OpenRouter tokens (or DeepSeek or MiMo or GLM or whatever), or rented GPU time.
The reality is that right now, there's nothing in the gap between 64GB and 128GB to make it worth a big upgrade cost. I seem to recall you have a 64GB Mac already? That runs Gemma 4 4-bit QAT, which I'm pretty sure is better than anything else that fits in 128GB. DS4 (the 1-bit quantization of DeepSeek V4 Flash) is too slow to be useful, and it doesn't feel smarter than Gemma 4 31B, to me..though I haven't used it enough to say that with confidence, since it's so slow.
@swelljoe
I’m currently using Deepseek and other Chinese models over OpenRouter.ai myself (at a cost of around $10/day) but I have friends that don’t have the option because of the sensitivity of their work. One just ordered a M5 Max MacBook Pro (won’t come until August) which will sit in his office running models - likely on some special cooling pad to make sure it keep everything running at peak…
My Mac mini M4 with 64GB has basically Gemma 4 12B on it most of the time for batch jobs of nonsense I do (one example: I finally achieved perfect email sorting by assigning definitions to folders and having Gemma 4 locally read my email and sort everything perfectly).
EDIT: want to be clear: as an indie dev, I work like 2-3 days a week so that $10/day figure is not as insane as it sounds.
@bryan 12B is the king of vision. Nothing self-hostable even on the Strix Halo beats it on vision tasks. Google is flexing with their open models.
But, yeah, I guess if you must self-host, and can't wait for normalcy to return, the best thing going is one of the 128GB Blackwell machines at $4000-$5000. I don't think any of the Strix Halo models are in stock, and since they've ramped up in price to be almost expensive as the Nvidia-based boxes, you might as well go Nvidia. That might change when the 495+ machines arrive with 192GB of RAM. (Though performance will be an issue on models that need that much RAM. As it is, any model that really fills up the 128GB on the Strix Halo is uncomfortably slow to use, and the 495+ is only a little faster.)
@swelljoe
The speed on the M4 Pro in my Mac mini is not great. Gemma 4 12B 8bit with 250k context is 17.3/sec. I wish I got the M4 Max Mac Studio w/64GB or 128GB.
Pro is 50% the speed as Max because slower memory bus. (Or not as wide?).
@bryan for speed, the 26B A4B model 4-bit QAT is where it's at. But, it's notably dumber than 31B. Smarter than 12B for code, but not vision. I still do most things in the cloud, though. My employers pays for the Claude $100 plan, and I don't know how to use it faster than it refreshes without shipping garbage, so it works out fine.
31B of Gemma 4 is too slow on the Strix Halo and just about usable on my desktop with two old AMD GPUs (Pro V620, which have about twice the memory bandwidth of the Strix Halo). Not fast, but in the teens. The DGX Spark boxes have memory bandwidth a little faster than the Strix Halo, but slower than my five year old data center GPUs...so, probably also not very fast on dense models in that size range. I think the ideal for this kind of 128GB unified memory hardware would be a 70B A10B parameter MoE model, but there isn't a current open model in that range. I keep hoping there'll be a double-sized Gemma 4 at some point. For now, 2x 32GB GPU beats 128GB.
