r/LocalAIServers 48m ago

Update — v2.0.0 is out, built directly from your feedback

Thumbnail
gallery
Upvotes
**•   Fixed the biggest issue**: high-VRAM systems (24GB+) now actually get 14B–35B model suggestions instead of tiny outdated ones  
**•   Physics-based speed estimation**: throughput is now calculated from real GPU/RAM memory bandwidth instead of rough size buckets  
**•   Multi-GPU support**: VRAM across multiple cards is now aggregated instead of only reading the primary GPU  
**•   MoE-aware**: proper handling for MoE models (like 35B-A3B) using active parameter count for speed estimates  
**•   RAM speed/channel detection**: DDR4/DDR5 speed now factored into hybrid CPU+GPU estimates  
**•   Fixed the layout bug**: disk cards and featured recommendations are now responsive, no more manual window stretching  
**•   Live sync improvements**: auto-syncs from Hugging Face on startup, shows “last synced X min ago”, filters for uncensored models and release date range  
**•   SHA256 + antivirus false-positive explanation** added to the README for anyone who got a scary flag

GitHub: https://github.com/keplerTR/LocalAI-Advisor

Thanks again to everyone who commented last time — several of these came directly from your suggestions (multi-GPU, MoE offload estimation, the layout fix). Keep the feedback coming.

Special thanks to u/Pika357, u/tbbtbbt, u/SnooOwls412, u/Jstratos9, u/Equivalent_Bass_879, u/QuarterDistinct857, u/arthax33, and u/RevolutionarySeven7 — several of the fixes above came directly from your comments on the original post.


r/LocalAIServers 2h ago

CMP170Hx “Spark” Machine

Thumbnail gallery
4 Upvotes

r/LocalAIServers 2h ago

Best LLMs for lower end hardware?

2 Upvotes

Im not sure if this belongs in here. But im trying to figure out the best LLMs to use on my 12gb Rtx 2060. I'm not looking for something that excels in one area or another but something general use that can handle extended conversations, light to medium co pilot coding, and general everyday use.


r/LocalAIServers 5h ago

What would you do with 1.5TB of ddr4 ram?

4 Upvotes

So, I have a lot of ram, 1.5 TB ddr4 ecc rdimm in 32gb sticks in old servers running a proxmox cluster, they don't support any decent gpus. Thinking maybe making an AI server for MoE models like glm5.3 flash UD-Q4_K_XL, and adding a gpu. Ideally I'd want a board with 16 ram slots and support for pciex16 3.0 or 4.0.

Or would it be better to sell all or most of it and try to get a more modern 512gb unified memory system?


r/LocalAIServers 9h ago

Open sourced our k8s-native AI platform for distributed multi-model inference at scale

2 Upvotes

Hey everyone,

I’m one of the co-founders of axem. We recently open sourced shaide, the infrastructure we’ve been building for running multiple LLMs across GPU servers.

It’s probably not very interesting if your setup is one model on one GPU. We started building it when our setups grew beyond that and we needed to run several models at the same time, scale them independently across GPU nodes, route traffic between replicas, and keep everything running inside infrastructure we controlled.

Instead of configuring all of those pieces separately for every cluster, we ended up packaging them into one Kubernetes-native platform.

Current setup:

  • vLLM for inference
  • llm-d for multi-instance orchestration
  • multiple models running and scaling independently
  • KV-cache-aware scheduling
  • internal OCI registry for container images + model weights
  • OpenAI-compatible API
  • the entire platform is managed as infrastructure as code
  • interactive installer that runs from Docker against an existing Kubernetes cluster
  • can operate fully air-gapped with no cluster egress

We mainly use it with on-prem RKE2 clusters, but it also works with EKS/GKE/AKS.

It’s Apache 2.0 and still early, so we’re at the point where feedback from people actually running multi-GPU servers is especially useful.

GitHub:
https://github.com/axem-solutions/shaide

I'm curious how people here handle this once a setup grows beyond a single machine.

If you're running several models in parallel, what are you using for routing/orchestration?


r/LocalAIServers 9h ago

GPT-OSS-20B quizá estaba muy infravalorado porque lo estábamos usando con el arnés equivocado

Thumbnail
1 Upvotes

r/LocalAIServers 11h ago

#005 32x CMP170HX 2TB VRAM Gen2.0@16x Deep diving into the numbers of speeds in system #rh3d

Thumbnail
youtube.com
5 Upvotes

r/LocalAIServers 17h ago

built from random parts

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/LocalAIServers 21h ago

Llm local run

Thumbnail
1 Upvotes

Can i run 32 gb laptop in run 27 b parameter model run?


r/LocalAIServers 21h ago

IA locale, quelle config ?

2 Upvotes

Hi,

I’m looking to run a local AI model on a laptop. The business use case involves managing highly complex building files, with several thousand documents forming a knowledge base.

I’m considering an on-premises laptop setup with an Intel i7, 64 GB of RAM, and an RTX 4090 with 16 GB of VRAM.

It should theoretically be able to run a 30B-parameter model, but I’m still looking for the best model to use.

Does this sound like a good approach?
Welcome.


r/LocalAIServers 22h ago

If you’ve set up local AI on Linux what actually broke, and how long did it take fix it?

4 Upvotes

I’m researching how developers set up vendor AI toolchains on Linux, such as ROCm/Ryzen AI, OpenVINO/oneAPI, CUDA, and others. I’d rather hear firsthand accounts than make assumptions.

1.Which stack, hardware, and distribution did you use?
2.How long did it take from a fresh installation to a model running on the accelerator?
3.What caused the issue? If you remember, provide specific details, such as a package, path, driver, or compiler version.
4. How did you verify that the GPU or NPU was being used and not silently falling back to the CPU?
5. Did you set up multiple vendor stacks? Was the second one easier to configure, or did most of the knowledge transfer not occur?
6. Did you document the process or create scripts to avoid repeating it?
7. If someone gave you a laptop with different silicon tomorrow and asked for the same setup, how would you feel about it?

I’m happy to share a summary of my findings with the thread.


r/LocalAIServers 1d ago

Will two RTX 3060 12GB cards be worth it for local LLM inference on a ThinkStation P520?

Thumbnail
5 Upvotes

r/LocalAIServers 1d ago

New setup on minis forum X1 pro-470 and R9700 Ai Pro 32gb dGPU over OCuLink

1 Upvotes

R9700 32GB eGPU + Minisforum X1 Pro-470: Qwen3.8-27B at 128k with image gen, and a counterintuitive Vulkan finding

Built this over the weekend and hit a few things that go against the usual advice, so figured I'd write it up.

\## Hardware

\- Minisforum AI X1 Pro-470 (Ryzen AI 9 HX 470, Radeon 890M iGPU)

\- 64GB DDR5-5600 (2x32, matched)

\- 2x Samsung 990 Pro 1TB

\- AMD Radeon AI PRO R9700 32GB in an AOOSTAR AG02 dock over \*OCuLink\*

\- Ubuntu Server 24.04.4, headless

\- llama.cpp + stable-diffusion.cpp, both built natively

\## The main finding: use Vulkan on the iGPU, not ROCm

This is the one I'd not seen mentioned anywhere. I wanted image generation running alongside the LLM, but a 27B model at Q4 plus Z-Image Turbo doesn't fit in 32GB together. Obvious answer: put image gen on the iGPU, which is otherwise idle.

With \*ROCm\* on the 890M: 2m09s for a 512x512 image, 8 steps.

With \*Vulkan\* on the same iGPU, same everything else: \*28 seconds\*.

4-5x faster. Vulkan reports \`uma: 1\` for the integrated GPU and appears to skip memory copies that ROCm makes on a device that shares system RAM. ROCm treats it like a discrete card.

Read lots online about "use ROCm on AMD" and that's correct for the R9700 — I use ROCm there. For integrated graphics it seems to that Vulkan is better, at least on RDNA 3.5. Maybe I missed something but for now this is working great for my useage.

Note the device-selection env vars aren't interchangeable: \`HIP_VISIBLE_DEVICES\` for ROCm builds, \`GGML_VK_VISIBLE_DEVICES\` for Vulkan. Setting the wrong one silently does nothing and you end up back on the discrete card wondering why it's fast.

Moving image generation to the iGPU freed up \~11GB on the R9700, which is what made 128k context possible.

\## Idle power: 92W -> 19W

llama-server was holding the card at full clocks doing nothing. \~90W+ constantly on an always-on box. No need for that nonsense.

Two flags fixed most of it:

\- \`--poll 0\` — stops the busy-wait. CPU went from 86% of a core to 0.6%.

\- \`--sleep-idle-seconds 60\` — unloads the model after a minute idle, releases VRAM.

That got VRAM freed but the card still sat at 3400MHz. Added a small script that polls VRAM usage and flips \`rocm-smi --setperflevel\` between \`low\` and \`auto\` depending on whether the model is resident.

Result: \*19W idle, 6.2s cold start\* on the first message of a session. Worth it for \~640 kWh/year.

Worth knowing: \`rocm-smi\` will still report GPU 100% while idle. Using amdgpu_top shows why — the command processor spins on an empty queue while every actual shader engine sits at 0%. It's a reporting artefact, power and clocks are the truth.

\## MTP speculative decoding is worth the effort

Unsloth ship an MTP module for Qwen3.8-27B in a separate \`MTP/\` folder in the GGUF repo. 1.3GB.

\--spec-type draft-mtp

\--spec-draft-model .../MTP/mtp-Qwen3.8-27B-Q4_0.gguf

\--spec-draft-n-max 3

Baseline tg128 without it: 24.8 t/s (Q5)

With MTP on real generations: \*38-52 t/s\* depending on workload - it made a huge difference to how it feels.

Draft acceptance runs 55-79%. Reasoning-heavy output accepts better than short answers — makes sense, it's more predictable. I tested n-max 2/3/4 and 3 was best for me, though the differences were a few percent.

\## Q4 vs Q6: Q4 still behaves on my workflows.

I assumed I'd want Q6. Ran both against a nasty multi-rule logic puzzle (nested conditional rules, some of which cancel others depending on question parity and primality). Same prompt, same output length:

| Q6_K_XL | Q4_K_XL |

| tg | 47.2 t/s | 50.4 t/s |

| pp | 799 t/s | 968 t/s |

| draft acceptance | 77% | 77% |

| VRAM @ 80k | 92% | 69% |

| answers | all correct | all correct |

Q4 is faster on both and uses 23 %points less VRAM. Unsloth's UD quants seem to hold up genuinely well on dense models. I'd previously seen bad hallucination from a \*\*MoE\*\* at Q4 — that's a different situation, each expert has fewer params so quantisation hits harder.

\## Final numbers

Qwen3.8-27B-UD-Q4_K_XL, 128k ctx, q8_0 KV, flash attention, MTP n-max 3, vision (mmproj), 2000 token reasoning budget:

\- pp512: 968 t/s

\- tg: 38-52 t/s in real use

\- VRAM: 75% of 32GB

\- Image gen (iGPU, Vulkan): \~28s per 512x512

\- Idle: 19W

Both models coexist. Web search, RAG, vision and image generation all work in one conversation. Access is Open WebUI behind \`tailscale serve\`, plus OpenCode on my laptop for coding work.

\## Other things that cost me time

\- \`nomodeset\` was needed to get through the Ubuntu installer\* on this hardware (console/framebuffer issue), and \*must be removed after\*, or amdgpu never loads and ROCm silently doesn't work.

\- The console goes dark during boot once amdgpu loads. as expected I guess. Use SSH.

\- ROCm needs the \*DKMS driver\*. Installing with \`--no-dkms\` leaves \`rocminfo\` reporting "ROCk module is NOT live" while everything looks superficially fine.

\- \`--parallel\` defaults to 4, which quadruples your KV cache. Set it to 1 if you're the only user. This was invisible to me for a while and I was blaming context size for VRAM pressure that wasn't context's fault.

\- \*OCuLink power limit\*: there are reports of AMD dGPUs being capped to the APU's TDP over OCuLink. Seems that's hat's a \*Windows driver\* issue but I never tested Windows so I can't confirm — on Linux mine draws the full 300W. Confirmed Gen4x4 link speed via \`amdgpu_top\`.

\- VAE decode needs a big compute buffer. \`--vae-tiling --vae-tile-size 16x16\` was the difference between working and OOM when things were tight.

Hope somebody gets some help from this.


r/LocalAIServers 1d ago

Where to get hardware?

Post image
20 Upvotes

I want to create this post to ask where everyone is getting their parts to put together their rigs. The only place I’ve been sourcing from is FB MKT and Ebay.

I’ve noticed the algo on both platforms shows what it wants me to see and hides postings. For example I’m suppose to believe there are only 10 people selling 2TB SSD NVMe Samsung 990 pro w/ Heatsink on Ebay but as soon I buy one there are 40 more selling it for $100 less than what I pay for?

Anyways just throw your recommendations/tips on finding hardware down below!!


r/LocalAIServers 1d ago

IRIS AGENT SYSTEM

0 Upvotes

🚀 Meet IRIS v0.2.0 – The Spatial Desktop Operating Environment for Autonomous AI Agents! 🧠💻

Most AI coding tools today are just single-stream chat boxes in a browser tab where you spend all day copy-pasting code snippets back and forth.

We decided to rethink how humans and autonomous agents collaborate. Meet IRIS (Intelligent Reasoning & Integration System).

IRIS isn't a chatbot. It’s a graphical agent operating environment built from scratch in Rust (Tauri 2) and React 19 / TypeScript. It treats agents, workspaces, tools, memory graphs, and release pipelines as first-class spatial desktop objects that you can arrange, inspect, run concurrently, and monitor in real time.

🔥 What’s New in v0.2.0:

🐙 1. GitHub Live Operations & Release Automation Connect your GitHub account in seconds. Specialist GitHub agents can triage open issues live, open surgical pull requests, automate SemVer releases (v0.2.0), author changelogs, and trigger GitHub Actions workflows that compile production binary builds (.AppImage, .dmg, .exe).

⚡ 2. Dual-Tier AI & Instant "Takeover" Stop overpaying for simple queries. Run fast, ultra-budget models (like Qwen 2.5 Coder, DeepSeek V3, or GPT-4o-mini) for 90% of routine workflows. When hitting a tough compiler error or tricky architectural refactoring, click ⚡ Takeover — a pre-configured heavyweight reasoning model (Claude 3.7 Sonnet, DeepSeek R1, Qwen 72B) immediately takes over the active conversation context with full reasoning depth!

🛸 3. Floating Desktop Desklet (Live HUD) Close the main window, and IRIS seamlessly condenses into a translucent, floating glass mini-HUD in the corner of your physical desktop. It displays real-time CPU/RAM telemetry, live agent thoughts, and keeps running smoothly as a background daemon.

🛡️ 4. Zero-Surprise Workspace Security & Visual Diff Viewer Inspect and approve exact code diffs before anything touches your local disk. All API keys and tokens are securely stored in your native OS Keyring.

🌟 100% Open Source (MIT License) & Local-First
Supports both local offline LLMs (via Ollama / vLLM) and all major cloud providers (OpenRouter, Anthropic, OpenAI, Google Gemini) plus standard Model Context Protocol (MCP) tools.

👉 Check out the repo, download the release, or drop a ⭐ on GitHub:
🔗 https://github.com/bubbadk/IRIS

I’d love to hear your thoughts: Do you prefer AI agents operating as spatial desktop applications rather than trapped inside browser chat tabs? Feedback and contributions are warmly welcome! 👇


r/LocalAIServers 1d ago

V100 or not?

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

Reliability of Tesla V100 SXM2 32GB on active cooled PCIe adapters

1 Upvotes

For a project I am looking at three or four Nvidia Tesla V100 32GB GPUs, but because of physical constraints, instead of buying the standard PCIe versions as usual, for this I am considering getting SXM2 V100 GPUs installed in those actively cooled PCIe-to-SXM2 adapters that are so plentiful on Aliexpress.

For the adapter I am eyeing something like this: https://www.aliexpress.com/item/1005012401568767.html

This seems to be the original version where the fan isn't temperature controlled and always runs full speed and where the slot cover has vents across the whole width and is held by screw at the backside of the PCB.

There is another version of that adapter where the fan is temperature controlled, this one has a small LED numeric display on the top and a slot cover which only has half of it covered in vents while the other half has the mounting screws:

https://www.aliexpress.com/item/1005011797591813

What worries me about the 2nd version is that the half-sided vents most likely will impact airflow of a GPU that can create up to 300W of heat, so I would assume the original version is probably a better choice.

However, before taking the plunge, I would be curious to hear from other users of SXM2 V100 GPUs in the both adapter variants, how well they work for them, if they had any stability or reliability issues, and how loud the uncontrolled fan is.

Edit: water-cooling is out of the question for various reasons, as are these multi-SXM2 boards which connect via OcuLink or MiniSAS to a PCIe adapter. It has to be a double slot PCIe card format and it has to have active cooling.


r/LocalAIServers 1d ago

WRX80 board stuck in shipping limbo for 3 months - TR PRO or switch to EPYC SP3?

3 Upvotes

Hey Everyone,

I'm currently building a local AI workstation/server. I've planned to base it on the Threadripper Pro 3000 Series (probably one of the 3945, 3955, or 3975 CPUs), mainly because DDR4 memory is way cheaper than DDR5.

In May, I ordered an Asus Pro WRX80 Sage from a German shop for around 1000 Euro (excluding tax), but it still hasn't been delivered. The delivery date keeps getting pushed back by 2-3 days every 2 days.

GPUs (2x R9700) and the PSU (Asus WS 2200W) have already arrived, and I want to start building =(

Therefore, I was looking for other sources. Unfortunately, no one else is offering the MB here in Germany. So, I'm left with eBay. On eBay, I found the following offers:

  • full box (box, MB, antenna, PCIe NVMe card) [China]: 1399€
  • full box (box, MB, antenna, PCIe NVMe card) [China]: 1499€
  • full box (box, MB, antenna, PCIe NVMe card) [Slovakia]: 1480€
  • only MB [UK]: 1250€

I'm no longer able to make offers to the Chinese sellers since eBay locked me out after they rejected my previous offers. I just found the Slovak listing a few minutes ago, so I haven't made an offer there yet.

While waiting for the initially ordered MB, I've been thinking about my choice for weeks now. Alternatively, instead of TR, I thought of EPYC SP3. My choices would be:

  • ASRock Rack ROMED8-2T [Finland]: ~950€
  • Supermicro H12SSL-i [China]: ~850€

I know EPYC CPUs are generally noticeably slower per-core than Threadripper, but they offer more cores and are cheaper. Still, TR is just really appealing to me.

What would you do - go with TR or Epyc, and why?


r/LocalAIServers 1d ago

I built a tool that scans your Windows PC and tells you which local AI models it can actually run

Post image
127 Upvotes

What it does:

• Scans your CPU (AVX2/AVX-512 support), GPU VRAM, RAM, and all your drives (NVMe/SSD/HDD)

• Pulls live model data from Hugging Face

• Tells you which models fit your hardware and at roughly what speed (tok/s)

• One-click copy for the ollama run command

It’s a standalone Windows .exe, no install needed. Fully open source (MIT license).

I’m not a professional developer — this genuinely started as a weekend experiment and grew into something I actually use daily. Would love feedback, bug reports, or feature ideas if anyone tries it out.

🔗 GitHub: https://github.com/keplerTR/LocalAI-Advisor


r/LocalAIServers 1d ago

Crazy FOMO - Can you help me please decide?

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

KeepRoLLMing v0.9.3 — an OpenAI-compatible proxy for more reliable local LLM chats and agents

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Are 24gb vram laptops sufficient for AI work?

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Hardware Advice: Single R9700 vs Dual-GPU R9700+RX 9070XT

Thumbnail
2 Upvotes

r/LocalAIServers 2d ago

It ain't much, but its mine. Local Ollama Test box with an i5 11400 and GTX1070

6 Upvotes

I've been wanting to play around with running Ollama locally.

I had an awful HP mini ITX box with the 11400 in it, but it was super janky and would never boot correctly. And the 300 watt ITX power supply wasn't safe to use with that video card.

I bought a used motherboard on ebay. Put the i5 11400 in it, a new Cooler (because HP put a 45 watt cooler on a 65 watt CPU). Then found it wouldn't fit in the case. So I had to go to microcenter this morning to get a case.

Got it all buttoned up. Now I'm running qwen3:30b and tied it to slack.

My children have managed to make it go insane, so I had claude add a code word to the setup so when I DM it the password from Eurotrip - "FLŰGGÅƏNK∂€ČHIŒβØL∫ÊN" it wipes the last 24 hours from its memory.


r/LocalAIServers 2d ago

Local or paying more?

8 Upvotes

I technically need a 5x account, and I've needed one for a month now. I've tried other Plus plans to get around that.

Right now, I'm at a point where I don't know whether to upgrade to a Pro account or invest 5k in a server and set up local AI models, filtered through at most a Plus account.

GPT thinks I can achieve 85-95% of the results I'm currently getting with Frontier models, and easily benefit from setting up that setup. I'm not entirely convinced; what do you think?

Thanks in advance.