r/opencode 3d ago

M1 Max 64GB Opencode + Qwen 3.8 27B + ??

Hi all, if you have an M-series Mac with 64GB, plus opencode 1.18.25 and Qwen 3.8 27B working successfully outputting high context for coding (30,000-120,000 tokens) can you share what local provider you’re going with? LMStudio, oMLX, llama.ccp etc

I’ve been having issues with LMStudio just timing out mid-response using Qwen 3.8 27B Q6_0 GGUF or taking over an hour to process each prompt request opencode makes before token generation using Qwen 3.8 27B Q6_0 MLX

Has anyone got a good high context, reliable solution going for Qwen 3.8 coding?

2 Upvotes

8 comments sorted by

1

u/Lyelinn 2d ago

You should ask this in local llm or local lama subs but my best bet is you don't have enough memory for that, remember that you need to change vram limit on macs, also consider olmx

1

u/raw-power 2d ago

Thank you, yes I’ve increased vram limit already to 57GB, is omlx less memory hungry than lmstudio?

1

u/swordofgiant 2d ago

LMStudio Bionic, download models.. Toggle the Local API Server. Add the base Url as OpenAI in other harness.

I am currently using the DeepSeek Harness with Local AI API server through LMStudio.

1

u/raw-power 2d ago

Running 3.8 27B? High context? On M1 Max 64gb?

1

u/arfung39 2d ago

I have an m5 max with 64gb, and I’m running OpenCode with oMLX, and a Qwen 3.8 27B oQ4e distill, with lightning MTP on, and it works great. If you leave thinking to xhigh, it does take a long time to come back, so I often turn it down to medium. Have done reasonably long (several hours on a run), coding projects with this set up. It gets quite slow when context used >60-70k tokens.

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/raw-power 2d ago

Thank you! I don’t swap when I use GGUF but it ends up timing out. MLX is fine until around 60,000 context just super slow prefil on each prompt opencode gives lmstudio but then after 60,000 it swaps and then it just crawls to 5 hours on prefil which is just unusable