r/opencode • u/raw-power • 3d ago
M1 Max 64GB Opencode + Qwen 3.8 27B + ??
Hi all, if you have an M-series Mac with 64GB, plus opencode 1.18.25 and Qwen 3.8 27B working successfully outputting high context for coding (30,000-120,000 tokens) can you share what local provider you’re going with? LMStudio, oMLX, llama.ccp etc
I’ve been having issues with LMStudio just timing out mid-response using Qwen 3.8 27B Q6_0 GGUF or taking over an hour to process each prompt request opencode makes before token generation using Qwen 3.8 27B Q6_0 MLX
Has anyone got a good high context, reliable solution going for Qwen 3.8 coding?
1
u/swordofgiant 2d ago
LMStudio Bionic, download models.. Toggle the Local API Server. Add the base Url as OpenAI in other harness.
I am currently using the DeepSeek Harness with Local AI API server through LMStudio.
1
1
u/arfung39 2d ago
I have an m5 max with 64gb, and I’m running OpenCode with oMLX, and a Qwen 3.8 27B oQ4e distill, with lightning MTP on, and it works great. If you leave thinking to xhigh, it does take a long time to come back, so I often turn it down to medium. Have done reasonably long (several hours on a run), coding projects with this set up. It gets quite slow when context used >60-70k tokens.
1
2d ago
[removed] — view removed comment
1
u/raw-power 2d ago
Thank you! I don’t swap when I use GGUF but it ends up timing out. MLX is fine until around 60,000 context just super slow prefil on each prompt opencode gives lmstudio but then after 60,000 it swaps and then it just crawls to 5 hours on prefil which is just unusable
1
u/Lyelinn 2d ago
You should ask this in local llm or local lama subs but my best bet is you don't have enough memory for that, remember that you need to change vram limit on macs, also consider olmx