r/SelfHostedAI Apr 17 '25

Do you have a big idea for a SelfhostedAI project? Submit a post describing it and a moderator will post it on the SelfhostedAI Wiki along with a link to your original post.

3 Upvotes

Visit the SelfhostedAI Wiki!


r/SelfHostedAI 8h ago

Second round of testing my utility. Not an Ad...

Enable HLS to view with audio, or disable this notification

1 Upvotes

Hello fam.

So, unfortunately, last time. There were some bugs but .. thats why we do UAT. Right ? 🤓

I wanted to show my latest round of testing.

I tried to show more system info this time.

Hopefully I did an "ok" job.

I'm looking forward to constructive suggestions and feedback.

Thank you for your time.

Have a wonderful day.

Will 👊🙏


r/SelfHostedAI 9h ago

LatticeVale Atomic — free installer/lifecycle manager for a Hermes Agent stack and related self hosted apps (Linux)

1 Upvotes

LaticeVale Atomic started as a separate evolution of my older LatticeVale project, but it is not just the Windows/WSL version copied over to Linux. I recently converted from Windows to Bazzite Linux OS and wanted to install LatticeVale locally. So I had AI help me figure that out.

The original project was built around Windows, PowerShell, WSL2, Docker, and the problems that come with coordinating software across both Windows and a Linux VM. LaticeVale Atomic was redesigned around a completely different environment: immutable/atomic Linux, rootless containers, user-level system integration, SELinux, and hosts where modifying the base operating system is intentionally discouraged.

The goal is still similar: make a complicated Hermes setup much easier to install, repair, update, recover, and manage. The way it accomplishes that is now very different.

Instead of relying on PowerShell and WSL, the Atomic version uses Bash, Python, rootless Podman, Compose, user systemd services, and XDG-compatible desktop integration. It is designed to keep the operating system itself as untouched as possible and place LaticeVale's managed state in user-owned locations.

The current stack can manage Hermes along with Matrix/Synapse, PostgreSQL, SearXNG, Valkey, QMD, Honcho, Redis, and local Ollama inference. It also includes hardware-aware resource limits, AMD GPU acceleration support where available, recovery snapshots, state-aware repairs, migration logic, audits, desktop Start/Shut Down controls, and protection against accidentally modifying unrelated Podman workloads. It uses terminal triggers instead of built in user selection during script run. See documentation on that.

Bazzite is currently the only platform I have tested. The project also includes adapters and compatibility logic for other atomic or immutable Linux systems, but those environments have not all received any real-world testing.

The current Bazzite migration has been tested on my own system with SELinux enforcing, rootless Podman, AMD GPU acceleration, an external Obsidian vault, existing Podman workloads, and a full 12-service Hermes stack.

That does not mean I can guarantee it will work perfectly on every machine. Atomic Linux distributions differ quite a bit in how they handle containers, host tools, permissions, user services, and immutable system boundaries. If someone finds a failure or a weird edge case, I would genuinely like to know about it.

This is still a hobby project, it is free, and anyone is welcome to inspect it, fork it, change it, or build on it. If you do experiment with it, I would be interested in hearing what works, what breaks, and what you improve.

LaticeVale itself does not bundle the third-party projects it manages. The installer retrieves or builds those components from their respective upstream sources.

I strongly recommend reading the included README, installer documentation, security notes, and migration information before running it. There is a lot going on internally, so using an AI model to inspect the repository or explain individual parts of the documentation is also a perfectly reasonable way to understand the project before installing it.

One final warning: local AI can be demanding. Ollama models can use a significant amount of RAM and VRAM, especially with larger models or context sizes. LaticeVale Atomic includes adaptive resource calculations and limits to reduce the chance of the stack overwhelming the host, but hardware still matters.

LaticeVale Atomic v15.2.3 AT6 is live:

https://github.com/winagainfinigin/LatticeVale-Atomic

Original Windows WSL version:

https://www.reddit.com/r/SelfHostedAI/comments/1vtpfdr/latticevale_free_installerlifecycle_manager_for_a/


r/SelfHostedAI 9h ago

Engineers running open-source LLMs in production: what is the hardest part today?

Thumbnail
1 Upvotes

r/SelfHostedAI 20h ago

The local AI challenge is no longer finding a good model. It is running it efficiently

Post image
1 Upvotes

r/SelfHostedAI 1d ago

I built a touch-first Ubuntu workstation you can carry in your pocket. Looking for alpha testers.

Thumbnail
1 Upvotes

r/SelfHostedAI 1d ago

IRIS agent system

7 Upvotes

🚀 Meet IRIS v0.2.0 – The Spatial Desktop Operating Environment for Autonomous AI Agents! 🧠💻

Most AI coding tools today are just single-stream chat boxes in a browser tab where you spend all day copy-pasting code snippets back and forth.

We decided to rethink how humans and autonomous agents collaborate. Meet IRIS (Intelligent Reasoning & Integration System).

IRIS isn't a chatbot. It’s a graphical agent operating environment built from scratch in Rust (Tauri 2) and React 19 / TypeScript. It treats agents, workspaces, tools, memory graphs, and release pipelines as first-class spatial desktop objects that you can arrange, inspect, run concurrently, and monitor in real time.

🔥 What’s New in v0.2.0:

🐙 1. GitHub Live Operations & Release Automation Connect your GitHub account in seconds. Specialist GitHub agents can triage open issues live, open surgical pull requests, automate SemVer releases (v0.2.0), author changelogs, and trigger GitHub Actions workflows that compile production binary builds (.AppImage, .dmg, .exe).

⚡ 2. Dual-Tier AI & Instant "Takeover" Stop overpaying for simple queries. Run fast, ultra-budget models (like Qwen 2.5 Coder, DeepSeek V3, or GPT-4o-mini) for 90% of routine workflows. When hitting a tough compiler error or tricky architectural refactoring, click ⚡ Takeover — a pre-configured heavyweight reasoning model (Claude 3.7 Sonnet, DeepSeek R1, Qwen 72B) immediately takes over the active conversation context with full reasoning depth!

🛸 3. Floating Desktop Desklet (Live HUD) Close the main window, and IRIS seamlessly condenses into a translucent, floating glass mini-HUD in the corner of your physical desktop. It displays real-time CPU/RAM telemetry, live agent thoughts, and keeps running smoothly as a background daemon.

🛡️ 4. Zero-Surprise Workspace Security & Visual Diff Viewer Inspect and approve exact code diffs before anything touches your local disk. All API keys and tokens are securely stored in your native OS Keyring.

🌟 100% Open Source (MIT License) & Local-First
Supports both local offline LLMs (via Ollama / vLLM) and all major cloud providers (OpenRouter, Anthropic, OpenAI, Google Gemini) plus standard Model Context Protocol (MCP) tools.

👉 Check out the repo, download the release, or drop a ⭐ on GitHub:
🔗 https://github.com/bubbadk/IRIS

I’d love to hear your thoughts: Do you prefer AI agents operating as spatial desktop applications rather than trapped inside browser chat tabs? Feedback and contributions are warmly welcome! 👇


r/SelfHostedAI 1d ago

Build an Apple Native AI Client - BYOK , 5 Surfaces , Connect ANY Harness

Enable HLS to view with audio, or disable this notification

4 Upvotes

So, you run your own local AI Harness. It's configured exactly to your needs. MCP, Skills, Capabilities, Context. You love the independent Harnesses such as Deepseek Harness, Pi, Aider or LiteLLM.

But how can you connect it to your Apple devices to access from anywhere? Your Watch, Mac or CarPlay.

Well, here is Conduck - the Apple native BYOK AI client.

Free and open source :-) .

It uses your Apple iCloud extensively and connects DIRECTLY via https to your own machine. Nobody in-between!

Check it either on https://conduck.com or GitHub https://github.com/GigaDuckAI/conduck


r/SelfHostedAI 1d ago

Agent Benchmark Exam

Thumbnail
1 Upvotes

r/SelfHostedAI 1d ago

How I Turned My Security Cameras Into an Automatic Bird Identification System with BirdNet-Go

Thumbnail
jasontucker.blog
1 Upvotes

r/SelfHostedAI 2d ago

cachegate — self-hosted LLM cache/router with a Docker one-liner and a built-in cost dashboard

2 Upvotes
cachegate: a self-hosted proxy for Anthropic/OpenAI that caches responses (exact + semantic) and routes to the cheapest healthy provider. One OpenAI-compatible endpoint, MIT licensed, no account or telemetry anywhere.

Docker:
docker run -p 4000:4000 --env-file .env ghcr.io/idebunk/cachegate:latest

Runs as non-root, ships a real HEALTHCHECK against GET /health (docker ps shows healthy/unhealthy), and degrades cleanly with caching disabled if you don't point REDIS_URL at anything - it won't fail to start over a missing optional dependency.

Also includes a small built-in cost dashboard (GET /dashboard) - KPI tiles, cost-over-time, cost-by-provider, a 7/14/30-day range picker - gated by the same bearer key everything else uses, no separate login system to stand up.

Config is one .env file. Refuses to start with an open /v1 endpoint unless you set an internal key (or explicitly opt into insecure local dev) - didn't want this to be secure only if you remember an extra step.

Honest gap: only Anthropic + OpenAI as providers right now, and the semantic cache doesn't scale past a few hundred cached entries per model (brute-force scan, documented in the README, not hidden).

Repo + full docs: https://github.com/iDebunk/cachegate

r/SelfHostedAI 1d ago

ForthMCP is meant to connect your local MCP servers to remote AI like Claude without port forwarding

Thumbnail
1 Upvotes

r/SelfHostedAI 2d ago

KeepRoLLMing v0.9.3 — an OpenAI-compatible proxy for more reliable local LLM chats and agents

Thumbnail
0 Upvotes

r/SelfHostedAI 2d ago

Using my new utility for secure AI. Not an Ad...

Enable HLS to view with audio, or disable this notification

0 Upvotes

I wanted to do a short demo of the utility I'm working on. I'm not a pro at video so, hopefully its not too terrible.

I created a cryptographically signed, secure tunnel for remote AI use. Basically, no one can talk to my AI Agent but me. Isn't that how it should be tho?

Now if i need to use my local AI on my pc.

No worries...

I hope the video isn't too bad.

I will do a better one.

Thank you for your time

Much love 🙏👊


r/SelfHostedAI 2d ago

Mac Mini / Studio

1 Upvotes

I saw the preorders this past week for Mac mini and studio. The builds I’d like are I the $3k and $5k range respectively. I know this is def a higher price than building my own rig for a self hosted AI machine.

I am curious what opinions are on a custom built rig with separated hardware vs the “unified memory” architecture of Apple silicon? I recently set up ollama on my 16g RAM m1 MacBook Pro from 2021 and please rly surprised at performance.


r/SelfHostedAI 2d ago

Our ad revenue read $0 for weeks and not one exception was ever raised

0 Upvotes

We run a free, keyless flight-search MCP server. It is free because sponsored results on the responses pay for the backend calls. For weeks the ad revenue read $0 CPM.

$0 is also what low traffic looks like. And bad fill rates. And a dozen other boring explanations, so nothing ever pointed at the code.

The ads SDK has a helper, register_result_widget, that attaches the revenue widget to a search result. Internally it does its work through asyncio.run, which is a no-op under an already running event loop. Our server is async. The widget had never attached. On any request. Ever. No exception, no log line, no warning.

The workaround was to stop using the helper and pass the widget mapping through the render call instead, which never touches the loop.

The free server is the ad-funded one; the paid sibling takes your own key and carries no ads. The general shape is worse than the bug: a revenue path with no error path is a special kind of fragile, because the failure mode of every bug in it is just a smaller number.

If you self-host something that earns, what actually raises when the earning stops?


r/SelfHostedAI 2d ago

Built a way to clone a GitHub repo and see it live-running on my phone in seconds — no laptop needed. Looking for alpha testers.

Thumbnail
1 Upvotes

I kept running into the same annoyance: I’d have an idea, or need to check something on a project, and I’d be nowhere near my dev machine. Remote desktop apps are miserable on a phone screen — tiny cursor, laggy, not built for touch at all.
So I built TouchWorkstation. It’s not remote desktop. It’s a touch-first interface to a real Linux machine you own — you clone a repo, it installs deps and starts the dev server on the actual machine, and you get a real live preview of your app right there on your phone. Then you can point an AI coding agent (Claude Code, Codex, etc.) at that exact project and actually keep working, not just look at it.
A few other things it does:
• Real persistent terminal (tmux-backed, survives you closing the app)
• Docker container management from your phone
• One-tap install for common dev tools
• Everything stays on your own hardware — no public exposure, reachable over your LAN or your own VPN, nothing routed through anyone else’s servers
It’s genuinely alpha — rough edges, actively building it, but the core loop (clone → live preview → agent) works and it’s the reason I built the thing in the first place.
If this is something you’d actually use, I’m looking for a small batch of alpha testers: https://touchworkstation.com
Happy to answer questions about how it works under the hood — it’s an Express + React app on a Debian package with systemd/nginx, nothing exotic.


r/SelfHostedAI 3d ago

Running PrivateGPT locally for secure document AI: my full setup guide

4 Upvotes

I wanted an AI that could answer questions from my personal docs without leaking data. PrivateGPT does exactly that, but the setup can be tricky. I documented everything—dependencies, model setup, Docker, and common errors—so you don't have to struggle. Perfect for self-hosters who value privacy.

https://interconnectd.com/blog/279/install-privategpt-secure-local-ai-for-your-documents-2026-guide/


r/SelfHostedAI 3d ago

Introducing Nexus: An Open-Source Enterprise Intelligence Framework for Production AI

3 Upvotes

The Veloxs AI team has released Nexus as an open-source Enterprise Intelligence Framework for building governed RAG, AI agents, semantic search, and intelligent workflows.

Core capabilities

  • Enterprise data, document, application, and event-stream ingestion
  • Processing, enrichment, classification, and identity resolution
  • Vector, lexical, hybrid, and graph-based retrieval
  • Governed RAG with citations and confidence signals
  • AI-agent orchestration, guardrails, and workflow automation
  • APIs and application-serving capabilities
  • Identity, role-based access, policies, and audit logging
  • Evaluation, performance, reliability, security, and cost monitoring

Install from PyPI

pip install veloxs-nexus

Nexus is designed to remain flexible across cloud providers, models, and data systems. The framework has been validated through 181 deterministic tests across eight test suites and currently supports Veloxs AI products, including Contexion.

GitHub: https://github.com/Veloxs-ai/nexus
PyPI: https://pypi.org/project/veloxs-nexus/
Contact: [contact@veloxs.ai](mailto:contact@veloxs.ai)

We welcome feedback and contributions from AI engineers, platform architects, researchers, and teams building production AI systems.

What capabilities do you consider essential for moving RAG and agentic systems from prototypes into governed production environments?


r/SelfHostedAI 3d ago

I got tired of bloated agent frameworks, so I wrote a local-first Rust runtime that gives LLMs real Linux permissions, persistent tmux sessions, and actual shell tools. Just updated v5.

Thumbnail
2 Upvotes

r/SelfHostedAI 3d ago

I wrote a complete guide to installing PrivateGPT for secure local document AI

2 Upvotes

PrivateGPT lets you run AI on your own documents without sending anything to the cloud. I got it working and wrote a step-by-step guide covering installation, configuration, and troubleshooting. If you care about privacy and want a self-hosted AI assistant for your files, this might save you hours.

https://interconnectd.com/blog/279/install-privategpt-secure-local-ai-for-your-documents-2026-guide/


r/SelfHostedAI 4d ago

Under 3 Seconds

3 Upvotes

After a lot of iteration, I finally got Christine’s latency consistently down to under 3 seconds using Warranted Retrieval.

That matters because Christine is not a cloud wrapper. She is laptop-bound, runs with no internet access, and has to operate within the actual limits of local hardware. Getting the response path down into a consistently usable range was a major milestone for me.

Now that the latency fight is finally in a much better place, it’s time to focus much harder on Christine’s training.

The next phase for me is less about shaving milliseconds and more about improving: - domain depth - retrieval quality - abstraction across domains - reasoning consistency - task usefulness under strict local constraints

Current laptop: - CPU: Intel Core Ultra 9 285H - RAM: 33.8 GB total physical memory - GPU 1: NVIDIA GeForce RTX 5050 Laptop GPU - GPU 2: Intel Arc 140T GPU - NPU: Intel AI Boost

I’m especially interested in what other people are doing with NPUs.

Are any of you actually using the NPU in a meaningful way for local/offline AI right now? If so: - what workloads are you pushing onto it - is it helping with latency, power efficiency, or always-on assistant behavior - are you using it for STT, routing, embeddings, background inference, or something else - and is it genuinely useful, or mostly just there in theory

Would like to hear from people building real local systems, especially laptop-bound ones.


r/SelfHostedAI 4d ago

I’ve been building Titans: local-first memory and durable execution infrastructure for AI agents

Thumbnail
1 Upvotes

r/SelfHostedAI 5d ago

My phone is now my personal ai assistant!

Enable HLS to view with audio, or disable this notification

144 Upvotes

Hi everyone! I bought a Titan 2 because I wanted a phone that was a tool, not a screen to scroll. Keyboard, buttons, a battery that doesn't quit. What I didn't expect was the opposite problem: after a few weeks the phone was more capable than anything I had to point it at. I was typing fast, into the same apps as everyone else.

So I built something for it. It's called Jenny. Its a personal AI agent that runs entirely on the phone, in an embedded Python runtime. It replaces the home screen, so pressing Home opens a conversation but can be also used as normal application. It remembers things, does work on a schedule while the screen is off, and writes its own little apps when I ask.

I made it free and open source hoping in community support. Even thanks means a lot for me 😄

It's been my daily driver for a month, and this is the video on my own titan 2. Small detail you'll appreciate: on a device with a real keyboard you can just start typing no tapping the input field first. I wrote that because of this phone.

I will be glad to answer to techincal and non techinal questions!

Download on github


r/SelfHostedAI 4d ago

OpenHands Docker setup for sovereign AI: lessons learned from my build

3 Upvotes

I spent a weekend getting OpenHands (formerly OpenDevin) running in Docker with a local LLM backend. It’s not as plug-and-play as some projects, but the result is a completely self-hosted AI coding assistant. I documented the exact steps, including GPU configuration, API keys for local models, and how to avoid common errors. If you’re building your own sovereign AI stack, this guide should help.

https://interconnectd.com/blog/278/opendevin-openhands-docker-setup-build-a-sovereign-ai/