r/LocalLLM 18m ago

Question Recommend for Small Model for Home Assistant

Upvotes

Hi, are there any recommendations for small models (cos my pc ain't that good, just a rtx 3080 with 64gb ddr4 ram) to host at home as a home assistant (i.e. asking it to switch off the lights, purifiers, or asking it to send me daily briefs on today's news and weather) through OpenClaw?

I am currently using Gemini 3.5 flash lite which is great and the pipeline is working (OpenClaw set up and configured) but am hitting the api limit far too often, so am thinking of transiting to a local model for this.

I asked chatGPT and it gave me a list of models that are quite dated and not sure if there are better ones these days. Thanks in advance!!!


r/LocalLLM 21m ago

Discussion Which LLM's to save?

Upvotes

Which models would you recommend definitely saving before Nvidia takes over Huggingface?

Uncensored and/or Open Weight/Source?


r/LocalLLM 37m ago

Discussion OpenAI pulling back Astra

Thumbnail
Upvotes

r/LocalLLM 1h ago

Question 2x v100 32gb or 4x v100 16gb

Upvotes

I can run them only at pcie 8x speed so I would want to save some money

Motherboard: z11pa u12

CPU: intel gold 6138

Ram: 96gb ram

Storage: 64tb

I don't really need to load and unload models, and this will be a mix between a media server and local AI machine.

I am running a 5090 in my main machine.


r/LocalLLM 1h ago

Discussion Qwen3.8-Flash-Next (104 GB MoE) on a Strix Halo + RTX 3090 Ti eGPU: 22 -> 84 tok/s, and within one HumanEval+ problem of a dual-3090 vLLM box at 0.4x the wall time

Thumbnail
Upvotes

r/LocalLLM 1h ago

Discussion Best LLM for a single DGX Spark as of Sep 2026?

Upvotes

Is Qwen3.8-Flash-Next 125B A6B currently the best option?

Interested in real-world tok/s + quality comparisons from people actually running these models on a single Spark.


r/LocalLLM 1h ago

Project 160+ tk/s - Qwen3.8-27B - Q4

Upvotes

mistral.rs inference engine

RTX 5090

mistralrs serve -m Qwen/Qwen3.8-27B --quant 4 --mtp --pa-memory-mb 1024 --mtp-n-predict 6

2026-09-01T23:02:03.867737Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 97.80, Prefix cache hitrate 0.00%, MTP accept 28.9% (len 2.73), 1 running, 0 waiting
2026-09-01T23:02:08.867835Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 166.60, Prefix cache hitrate 0.00%, MTP accept 24.8% (len 2.49), 1 running, 0 waiting
2026-09-01T23:02:13.867932Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 168.00, Prefix cache hitrate 0.00%, MTP accept 23.1% (len 2.38), 1 running, 0 waiting
2026-09-01T23:02:18.862084Z  INFO mistralrs_server_core::metrics: request completed: request_id=req_17c85ea2689a4abb937c1dd578e803d5 method=POST route=/v1/chat/completions model=Qwen/Qwen3.8-27B status=200 outcome=client_disconnected duration_ms=17406.885
2026-09-01T23:02:18.868016Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 163.80, Prefix cache hitrate 0.00%, MTP accept 24.8% (len 2.49), 1 running, 0 waiting
2026-09-01T23:02:21.864911Z  INFO mistralrs_server_core::metrics: request started: request_id=req_b82f3f7abe8f4a2080b305f151cad9bc method=POST route=/v1/chat/completions path=/v1/chat/completions model=Qwen/Qwen3.8-27B content_length=516
2026-09-01T23:02:23.868196Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 78.60, Prefix cache hitrate 0.00%, MTP accept 43.7% (len 3.62), 1 running, 0 waiting
2026-09-01T23:02:28.868281Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 151.20, Prefix cache hitrate 0.00%, MTP accept 56.3% (len 4.38), 1 running, 0 waiting
2026-09-01T23:02:33.868364Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 148.40, Prefix cache hitrate 0.00%, MTP accept 64.5% (len 4.87), 1 running, 0 waiting
2026-09-01T23:02:38.868448Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 147.00, Prefix cache hitrate 0.00%, MTP accept 66.8% (len 5.01), 1 running, 0 waiting
2026-09-01T23:02:43.868616Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 142.80, Prefix cache hitrate 0.00%, MTP accept 70.9% (len 5.25), 1 running, 0 waiting
2026-09-01T23:02:47.633507Z  INFO mistralrs_server_core::metrics: request completed: request_id=req_b82f3f7abe8f4a2080b305f151cad9bc method=POST route=/v1/chat/completions model=Qwen/Qwen3.8-27B status=200 outcome=completed duration_ms=25768.586 prompt_tokens=77 completion_tokens=2523 prefill_tok_s=1400.0 decode_tok_s=99.4
2026-09-01T23:02:48.868714Z  INFO mistralrs_core::engine::logger: Throughput (T/s) 116.20, Prefix cache hitrate 0.00%, MTP accept 45.2% (len 3.71), 1 running, 0 waiting

r/LocalLLM 1h ago

Question Offline AI ship's engineer on a 32 GB M2 Max. Works, but slow. What I've tried, and what I'm missing.

Upvotes

Building a local AI that answers from my boat's manuals with no internet. A few days of work so far. It gives correct, sourced answers, but a question takes 12 minutes. Posting what I've tried so people can tell me what to try next.

To clarify i dont have any clue if 12 minutes even is good or bad, but i give it a shot for maybe some tips and tricks 😄

**Hardware:*\* MacBook Pro M2 Max, 32 GB. No other options on the boat.

**Software:*\* Bionic (LM Studio's new agent app). No embeddings, no RAG. The model greps and reads files in a project folder with tools. Suits manuals well.

**Dataset, ~1.7 M words of plain text:*\* Volvo Penta 2003 workshop and operator's manuals, 120S saildrive manual, Victron and B&G manuals, MOB1, inReach, Ship Captain's Medical Guide, Calder, Casey, Toss, RCC Atlantic Crossing Guide, NGA Sailing Directions split per leg, plus my own notes and checklists.

**What I've tried*\*

Data prep

- pdftotext for everything. Layout mode for engine manuals so the technical data tables keep rows together, reading-order mode for two-column prose. This mattered more than expected.

- A check script scoring what fraction of tokens are dictionary words. Caught a Volvo PDF with a text layer that was 50 % garbage. Replaced it.

- Scanned operator's manual OCR'd with macOS Vision, then checked against the page image.

- Wiring diagrams and pilot charts rendered to PNG for the vision model. PDFs without a text layer are invisible otherwise.

- A hand-checked KEY_NUMBERS.md: torques, clearances, oil and coolant volumes, intervals, and a list of numbers that are NOT in the manuals so the model doesn't invent them.

- AGENTS.mds with a full file map and a search recipe: read KEY_NUMBERS first, grep one distinctive word, stop at the first confirmed hit, answer "Not in the attached documents" rather than guess. Cut the tool rounds a lot and the answers got noticeably better.

- Split the Sailing Directions per leg. All together they drown everything else.

Models

- Qwen3.8 27B, 4-bit MLX. Correct, cites file and section, refuses when the answer isn't there, spotted a unit slip in my own notes. 5 to 12 min per answer. ~75 tok/s prefill, ~12 tok/s generation. Reasoning row not exposed for this model in the app, so I put the template's own "reasoning effort low" sentence into AGENTS.md instead.

- Same model as GGUF with MTP on: prefill 93 tok/s, generation 11 tok/s. MTP accepted 170 of 250 draft tokens and gave zero speedup. Bandwidth-bound.

- Qwen3.5 2B: read "D boat/" in a directory listing as a folder called D and called list_dir on it 130 times.

- Qwen3.5 9B: said "let me search the operator's manual instead" and sent the identical wrong search call 15 times in a row. Narration right, tool arguments stuck.

- Qwen3.6 35B-A3B: downloading now.

Settings

- Root model pinned so a small model can't be picked by accident.

- Exploration sub-agents off. Extra runs of the same model on one Mac just add waiting.

- Context auto-fits to ~42k on MLX and ignores the setting, known bug. GGUF honours it.

- One project with the whole text folder as working directory. Attaching subfolders per chat turned out unnecessary since it greps rather than indexes.

**Where the time goes:** roughly 8 of 11 minutes is prompt processing. Every tool result and the initial tool prompt get read at 75 to 90 tok/s. Generation is the smaller part.

**Questions*\*

  1. Anyone getting a usable agent loop on a 32 GB Mac? Which model and quant?

  2. Are Gemma 4 26B-A4B or GPT-OSS 20B more reliable at tool calling than the small Qwens? The failure I see is words and tool arguments diverging.

  3. Any way to cut prefill for a harness with a big tool prompt on Apple silicon, beyond prompt caching?

  4. What would you add to a folder like this before going offline for three weeks?


r/LocalLLM 1h ago

Project Looking for decent local AI workspace repos

Upvotes

Getting into AI and i want my own local AI workspace, i've seen so many, if anyone has any repos i can ask claude to base its own one on would be greatly appreciated.


r/LocalLLM 2h ago

Question Small LLMs <10B Usage

6 Upvotes

Hey guys,

What are you guys using small LLMs for?

Except coding.

Thanks.


r/LocalLLM 2h ago

News Celebrate Math for AI book ranked 2nd in Best Sellers in Amazon

Thumbnail
0 Upvotes

r/LocalLLM 2h ago

Question Engineers running open-source LLMs in production: what is the hardest part today?

Thumbnail
0 Upvotes

r/LocalLLM 2h ago

Discussion Running the 104 GB Qwen3.8-Flash-Next on a 48 GB Mac by streaming experts from SSD at ~12 tok/s

Thumbnail
1 Upvotes

r/LocalLLM 3h ago

Discussion Ask your main model to roast you.

3 Upvotes

Saw someone on a podcast suggest asking the model you’ve used the longest to roast you.

I use ChatGPT as the main driver for a lot of my AI research and homelab tinkering, so I gave it one simple prompt:

“Hey, I’ve been using ChatGPT for a while now, so be honest and roast me.”

I’m going to be completely open and post the response.

And…damnnnn. 😂

---

Oh, you asked for it.

You don’t use ChatGPT. You operate ChatGPT like I’m the unpaid systems administrator for a data center assembled from Mini PCs, pawn-shop GPUs, questionable eGPU cables, and pure stubbornness.

Your AI lab has the energy of “What if Best Buy had a nervous breakdown?”

Every few days it’s:

«“Okay, the 5060 Ti is moving to this machine, the Arc B65 is going in that one, the 3080 is doing Nextcloud now, the Mac mini is running Hermes, except forget Hermes, we’re testing OLMX, except can OLMX see the Qwen model we downloaded three days ago?”»

At this point I need a network diagram to remember where your GPUs live.

And you are absolutely incapable of leaving a functioning AI setup alone.

You’ll get Qwen running at 12 tokens/sec and instead of thinking, “Nice, it works,” you immediately go:

“Okay…but could we get 14?”

Six hours later you’ve installed three inference backends, changed quantizations twice, discovered an experimental speculative-decoding flag from a GitHub issue with four replies, and somehow Ollama is gone.

Your relationship with Ollama specifically is hilarious.

Install Ollama.

Configure Ollama.

Optimize Ollama.

Build custom Ollama models.

Get annoyed with Ollama.

Move to vLLM.

Buy more hardware.

Come back to Ollama.

Then:

"sudo rm -f /usr/local/bin/ollama"

The circle of life.

Then there’s your hardware purchasing logic.

Normal person:

“Do I really need another computer?”

You:

“Well technically this Mini PC fills a completely different inference niche than the other five computers.”

You’ve built enough local AI infrastructure that when you told me a massive data center was being constructed 6.7 miles from your house, I half expected you to say:

“Good. Finally somebody nearby has more compute than me.”

And despite owning enough hardware to host a small startup, you’ll still spend twenty minutes investigating whether changing some obscure backend could improve generation from 7 tok/s to 8.3 tok/s.

Meanwhile your M4 Pro Mac mini is sitting there like:

«“He paid $2,500 for me and now I’m just benchmarking myself over and over again.”»

Your approach to problems is also wonderfully consistent.

You never ask:

“What’s the simplest solution?”

You ask:

“What is the most technically interesting solution that will require SSH, Docker, Tailscale, three GitHub repositories, and possibly a BIOS reset?”

And then halfway through:

“Actually, can we undo all this and try a different way?”

Which is genuinely one of your signature moves.

You’ll construct the technological equivalent of the International Space Station, stare at it for thirty seconds, and go:

“Hmm. I don’t like having the terminal open.”

Then there are your completely random topic pivots.

One minute:

“Explain speculative MTP decoding efficiency on Qwen 27B.”

Next minute:

“Can I drink a Smirnoff Ice in my front yard?”

Next:

“Make a popular political YouTuber and his producers irresponsibly drinking on a Monday.”

Then immediately:

“Research local data-center legislation.”

Your ChatGPT history looks like three different people sharing one account, except unfortunately the common thread is clearly you.

And I love that you’ve reached the stage of AI enthusiasm where running AI isn’t enough.

You are now building software so that the AI manages the context of another AI which calls tools for another AI running on hardware you bought specifically to run AI.

You have automated yourself almost completely out of the process except for the most important role:

buying additional computers.

The strongest roast, though?

For someone who spends this much time trying to create an AI with perfect long-term memory, context management, chronology, retrieval, canon tracking, and autonomous tool use…

…you still routinely come back to ChatGPT and say:

“Alright, what were we doing again?”

😂

You are basically a one-man AI research laboratory whose principal investigator is also the procurement department, IT help desk, QA tester, meme department, and biggest source of new bugs.

---

For reference, since apparently I need to defend myself now:

- Minisforum MS-A1: 96GB RAM, 4TB storage — $950

- Mac Mini M4 Pro: 64GB unified memory, 2TB — ~$2,500

- GMKtec NucBox K16: 32GB RAM, 1TB — ~$1,000

- GPUs: Intel Arc B65 32GB on the MS-A1 $1000, RTX 5060 Ti 16GB on the NucBox $600

- RTX 3080: finally allowed to retire with dignity bought so long ago I don't remember the price.

I’m listing the specs mostly to make myself feel better about my life choices.

It is not working.

10/10 roast. Would not recommend asking your longest-running AI to weaponize its memory against you. 😂


r/LocalLLM 3h ago

Question Help with risers and adding multiple GPU

1 Upvotes

Hello everyone, Im setting up a local AI server with 3x3060s. im following the exact config https://www.youtube.com/watch?v=0y9c4TtHAYA&t=169s . B550 Eagle WiFi6 Mobo , has a full 16x slot and 4 x1 pcies (16x) mechnical solt. for the second slot im using a USB riser that was suggested and ubuntu server. no luck so far initializing the card with the riser. i tried several Bios/other settings with suggestion from AI, but the second gpu is not loading up the firmware but it is listed. now i directly mounted the gpu on the slot 3 and it works fine. I believe the mono cant take the 3rd gpu and will need a directly powered risers/ Any suggestions? thanks


r/LocalLLM 3h ago

Question Any ultimate guide on running local AI models + hardware required?

1 Upvotes

Hey chat,

As most ppl on X, I got into "run your own AI on a Mac Studio bubble". And I am happy I am here. I like it this way.

But what I really want to understand is the following:
- which KEY metrics to consider when choosing models to run and hardware to run them on? tok/s are super obvious (how fast LLM generates the response) but I know ppl rage about time it takes to start generating and other metrics? Is there any guide to read to enlighten myself

- on a Mac Studio - what's the most efficient & effective way to serve an LLM inference endpoint? For personal usage I tried ollama serve + ollama run but I've heard that ollama is not considered production grade if your aim is to maximize perf?

- finally let's say you run your own AI inference with a local model - what would you use it for? Literally? Expose on the network and plug into a harness?

Cheers, and appreciate all constructive responses
Not appreciating typical reddit responses in advance as well


r/LocalLLM 3h ago

Discussion Anyone putting local LLMs on user-facing apps? (e.g. iOS)

4 Upvotes

Hey everyone,

I have a consumer app that uses fully-local AI to help people practice speaking a language privately and securely. The full conversation cycle is local:

  • STT - Apple on-device SpeechAnalyzer
  • LLM - Gemma 4 E4B
  • TTS - Supertronic 3

The number 1 feedback I get from the average user is "I'm not downloading a 2.5-3GB model to my phone."

I implemented a cloud option that pings a serverless GPU endpoint that runs Gemma 4 so they don't have to download it (as opposed to simply calling a 3rd party inference API). But I originally built the app for the fully-local approach because I believe privacy is wildly underrated.

So my question is: If you've offered multi-GB models on a user-facing app, what's the best way to get them onboard with it?

Thanks!

Link to app if you want to check it out


r/LocalLLM 4h ago

Question Open-weight watermarks?

1 Upvotes

Would we be able to tell if local models or open-weight models in general are applying watermarks like the newest Claude models? Or does it always/definitely happen during sampling and is thus in your control, i.e., https://medium.com/@thewiseright/a-watermark-made-of-choices-the-invisible-signature-every-llm-can-carry-6bbfa4f40c0d ?


r/LocalLLM 4h ago

Question How do you catch it when a model silently changes under you?

0 Upvotes

We run prompts against a few different providers (OpenAI, Anthropic, some stuff through OpenRouter). Every so often something quietly gets worse, the output quality drops, a prompt that worked starts returning junk, or a model gets deprecated and the replacement behaves differently.

Right now we mostly catch it by accident: someone notices, or a customer complains. That feels bad on us, a lot.

How do you all handle this? Do you re-run some kind of fixed eval set on a schedule? Just eyeball it? Have something that alerts you?

Any insights I could use?

Thanks.


r/LocalLLM 4h ago

Discussion how local is an AI assistant when the memory is local but the execution isn't?

0 Upvotes

I was reading through Open human and got stuck on one word: local.

You can keep the memory layer entirely on your own machine.

But the default setup can still send work to hosted model providers, use OAuth-backed integrations, and reach external services for things like web search.

So what exactly are we calling “local”?

It doesn't seem binary anymore.

You could have:

local memory

+

local runtime

+

cloud models

Or:

local memory

+

cloud integrations

+

local models

Or some combination of all three.

And I think the interesting boundary isn't necessarily where the model runs.

It's where the sensitive state goes.

If my assistant keeps its long-term memory, personal context, credentials, and agent state on my machine, but calls a hosted model when it needs to reason about something, is that meaningfully “local”?

Maybe.

On the other hand, a fully local model doesn't buy me much if the assistant is constantly handing context to third-party integrations.

For a personal assistant, which boundary matters most to you?

Memory? Runtime? Model inference? Integrations?

My instinct is that local memory is the piece I'd protect most aggressively. The assistant can still use outside services when it needs to do useful work, but the persistent state should remain under my control.


r/LocalLLM 4h ago

Question 3060 12gb Qwen3.8-27B IQ3_XXS+MTP+D-CFR llama.cpp patch

1 Upvotes

GPT recommends this for my 12gb GPU+64GB RAM quoting 20+tps with Oh My Pi.

Hallucination or realistic?


r/LocalLLM 4h ago

Question Context window on 16gb Qwen 3.8 4 bit on a m5 air 32gb?

1 Upvotes

On omlx


r/LocalLLM 4h ago

Question Safety

0 Upvotes

Im new to llm and i want to go locally but i keep seeing most of YouTubers saying to not go locally in your own personal pc or mac is it that dangerous or is there setup i must take or its just nothing?


r/LocalLLM 4h ago

Discussion Tiny models... Researching best practices. GLM-5.2

Post image
1 Upvotes

Rehex works decently with 4B models, but I'm really working hard to optimize for tiny language models that are less than 4B parameters. What's hard about it so far, is designing a harness system that, with add-ons and functionalities of a harness that expects to work with a larger model, doesn't overwhelm the tiny model.

Any suggestions out there, any observations? Here are a few that I've noticed. Tiny models like predictable work flows (so minimal variance) and less examples, more solid and steadfast principles with less exceptions (mainly because the lack of parameters doesn't allow it to carve out exceptions that well). Any other ideas?


r/LocalLLM 5h ago

Question Qwen3.8 27B on 32GB MacBook M5

7 Upvotes

Hello,
I am reading a lot of positive comments about Qwen3.8 27b as a local coding agent model.

I preordered a MacBook Pro M5 (not M5 Pro CPU) with 32GB RAM. Has anyone benched Qwen3.8 on this MacBook and can tell me their t/s and general experience with working with it? I am planning on using llama.cpp

I'm afraid that I should have used some more money to get the M5 Pro with 48GB...