r/LocalAIServers • u/Fast_Bar_1234 • 9h ago
Best LLMs for lower end hardware?
Im not sure if this belongs in here. But im trying to figure out the best LLMs to use on my 12gb Rtx 2060. I'm not looking for something that excels in one area or another but something general use that can handle extended conversations, light to medium co pilot coding, and general everyday use.
2
u/Proper-Tower2016 9h ago
probably Ornith 1.5 35b3a apex i-quality (at least if you have 32gb of RAM or more)
1
u/Glad_Contest_8014 7h ago
Need more info. Cpu? Ram? You are likely looking at a cpu offloaded model in all honesty. You can at least load weights to the vRAM to reduce I/O thrashing, which puts you in a reasonable 14tok/sec level with anything that fits on your RAM footprint.
1
u/simos_sayz 7h ago
"can handle extended conversations, light to medium co pilot coding, and general everyday use."
Gemma4:12b Q4 QAT would do the job for the extended conversation and general everyday use. Coding probably not that great. Just make sure you set up web search
1
u/Dorkits 9h ago
With this card? None.
Just get DeepSeek API and be happy.
2
u/Unnamed-3891 8h ago
”None” 🤣
Qwen-3.6-35B-A3B will work just fine. And a miriad of much worse models. OP will obviously need to temper their expectations for speed and context size, but it will work.
1
u/searchblox_searchai 9h ago
You can run small efficient models for free locally https://inference-server.searchblox.com/index.html#install