Thread of 2 posts

jump to replies

I shouldn’t be distracted with this kind of nonsense when I have plenty of apps needing some tlc to get out the door… that said…

I’ve been spending time trying to get rapid-mlx to work on the Mac mini. It’s supposed to be much faster than llama.cpp, but I’ve had nothing but problems. Hope it is just a learning curve on the right mix of parameters. Defaults definitely don’t seem to be dialed in like llama-server. #opencode #rapidmlx #llamacpp

Using the exact same model, qwen3.6-35B-8bit, I get infinite looping issues with rapid vs llama. It’s like the conversion from gguf to mlx format breaks something. Could also be different default temps and other settings. 😑
#opencode #rapidmlx #llamacpp

2 replies

back to top
Santiago, né ? :amiga: 👾 , @santi@gone.lema.org
(open profile)

@bryan Saw your post and I installed rapid-mlx on my m1 studio and tested simple chat via open webui with the recommended fastest model (nemotron 30-B) and I got a loop quickly. Also I never mentioned anything about cannibalism … just ask things like how fast are you.

BTW this model is the most sycophantic I ever saw. To a point it’s comical.