Thread of 2 posts
jump to repliesI shouldn’t be distracted with this kind of nonsense when I have plenty of apps needing some tlc to get out the door… that said…
I’ve been spending time trying to get rapid-mlx to work on the Mac mini. It’s supposed to be much faster than llama.cpp, but I’ve had nothing but problems. Hope it is just a learning curve on the right mix of parameters. Defaults definitely don’t seem to be dialed in like llama-server. #opencode #rapidmlx #llamacpp
3 replies
back to top@bryan Saw your post and I installed rapid-mlx on my m1 studio and tested simple chat via open webui with the recommended fastest model (nemotron 30-B) and I got a loop quickly. Also I never mentioned anything about cannibalism … just ask things like how fast are you.
BTW this model is the most sycophantic I ever saw. To a point it’s comical.
@santi It is pretty weird. The review I was looking at gave it high marks. But I've just not had good results with any of the mlx based #LLM.
Example review: https://andrew.ooo/posts/rapid-mlx-fastest-apple-silicon-llm-server/
@santi Forgot to mention, nemotron has always been pretty trashy for me even in llama.cpp. It will get into weird loops even with curated / recommended model on LM Studio.

