r/AI_Agents 6d ago

Weekly Thread: Project Display

8 Upvotes

Weekly thread to show off your AI Agents and LLM Apps! Top voted projects will be featured in our weekly newsletter.


r/AI_Agents 1d ago

Weekly Hiring Thread

2 Upvotes

If you're hiring use this thread.

Include:

  1. Company Name
  2. Role Name
  3. Full Time/Part Time/Contract
  4. Role Description
  5. Salary Range
  6. Remote or Not
  7. Visa Sponsorship or Not

r/AI_Agents 1h ago

Discussion Is Cresta a good AI agent?

Upvotes

We’re a pretty big team and looking to implement an AI agent for a lot of the repetitive work in our contact center. This is one of the tools we’ve been looking at and so far most of the reviews seem pretty good.

Im interested on how it handles real contact center workloads and whether the setup is worth it. Just wanted to get some opinions from people who have used it before we go any further.


r/AI_Agents 6h ago

Tutorial Senior AI engineering interviews aren't definition questions. They're "your system just broke in prod, talk me through it."

22 Upvotes

Been going through AI/LLM interview prep material for a while and most of it plateaus early. Define RAG. Define an embedding. Explain prompt engineering. Fine for a first screen, useless for anything senior.

The interviews I've actually sat in — on both sides of the table — look nothing like that. They look like:

  • Accuracy is fine in staging, drops in production. Where do you start looking?
  • Latency went from 2s to 8s overnight, nothing in the code changed.
  • Spend went up 4× and nobody can say why.
  • Your agent is in an infinite tool loop.
  • Four teams have independently built four RAG platforms and now you own all of them.
  • Frontier model vs. small model vs. fine-tuned — defend the pick with numbers.
  • Security says customer data can't leave the VPC. Redesign.

And they don't stop at your first answer. You design the thing, then it's "traffic is 15× now," then "costs tripled," then "the provider is having an outage," then "accuracy is down 15%." The point isn't the answer, it's whether your reasoning survives the constraints changing under you.

Coding rounds are the same story. Less LeetCode, more: write an LLM client with retry/backoff, build async inference with cancellation, implement rate limiting, build an eval harness, write an agent loop that terminates. Happy path is table stakes. What they're watching for is timeouts, backpressure, failure modes, observability.

So I started writing all of this down as scenarios with worked reasoning rather than answer keys — open source, still rough in places.

Checkout link in the comments


r/AI_Agents 5h ago

Discussion There are 3,749 AI-run news sites now and most of them aren't written for humans at all.

14 Upvotes

Yea just imagine what if they start scraping each other, or chatgpt gets trained on the same slip that it generated. I was reading pratham mittal’s newsletter.

there is a platform that apparently tracks fully automated news sites and the count is past 3,749 now, across 16 languages. the business model really got me. a lot of these aren't chasing human readers at all. They publish so they get picked up by aggregators and other bots, and the ad impressions come off machine traffic.

Closest analogy i can think of is stale cache propagating through layers because nobody set an invalidation strategy. except the origin here is a human who wrote a thing once and then moved on with their life.

Anyway, the newsletter framed it as a content problem, but i think it's an infra problem. so asking people who deal with this properly: is there anything technical solution that can push back?


r/AI_Agents 11h ago

Discussion We sent a meme deck to a $400M company as a joke. They replied. It's our entire outbound now.

32 Upvotes

Ok so this started as a joke experiment and then it worked, so now it's a process.

Quick context so you know where this is coming from. We're an AI agency. The whole pitch is we find one expensive, repetitive problem inside a company and put AI around it so they don't have to hire a full time AI team. Which means our own outbound has to work or we don't eat. No SDRs and no outbound agency, we DM people ourselves. This is the workflow we use to get our clients.

Last month one of these got an actual reply, with a question in it, from a company doing about $400M. That is not a company that needs to reply to a small shop. We'd been sending them as a casual test and after that reply we made it the default.

I'm aware this reads like the exact kind of post I usually roll my eyes at. You can go do it in an hour and see for yourself though.

The core idea is dumb simple. Buying is psychological. You can stack 30 tools on top of a message that's aimed at the wrong person with the wrong framing and it still dies in the inbox. So before any writing happens we make Claude work out a few things about the decision maker that almost everyone skips:

How old are they. A 28 year old head of ops and a 58 year old COO do not respond to the same message and pretending otherwise is how you get ignored by both.

Who actually owns the budget. In a 50 person company the person who feels the expensive problem every day is usually a director and the C-suite just signs. Everyone DMs the CEO anyway.

How long they've been in the seat. Six months means they're still building the playbook and still shopping. Four years means they're locked in and you'll have to pry them out.

What's their stack. If they're still running 2021 software, the repetitive problem is probably sitting right on top of it.

What's changed lately. They posted a job for an in-house AI hire or a new ops exec came in or a launch is about to bury some team in manual work. Any of that means the problem is getting expensive right now.

What are they posting about. If a COO is suddenly posting about AI adoption, someone above them asked what the plan is and you want to be in the inbox that week.

We tuned all of this for what we sell. Swap the signals for yours.

Setup is small. Claude with the Gamma connector switched on, plus a LinkedIn account. No agent framework and no 30 tool stack. Everything below is just what we ask Claude to do, in order.

Step 1. Don't buy anything. LinkedIn settings > data privacy > get a copy of your data > tick connections. They email you a CSV in a day or two. Your existing network is a better list than anything you'd scrape and you already have a reason to message them.

Step 2. Find what the buyer actually wants. Paste your offer in one paragraph plus who you're targeting (title and company size, industry if it matters) and ask Claude for the real problem underneath the surface complaints and why they'd buy this week instead of next quarter. Push it past the obvious answers, the first pass is always generic. This is the raw material for everything after.

Step 3. Turn that into a person. Ask it to build a full persona from the same targeting. The words they use. What they're scared of at work. What they actually care about versus what they post. You want it to read like a human instead of a job title.

Step 4. Get your openers. Give it the offer plus the persona and ask for 8 different ways to open a conversation, each one built off something specific in the persona. None of them should be the "quick question" template every SDR on earth is running right now. If one of them is, throw it out.

Step 5. The meme deck. This is the part that gets replies. Ask Claude to look up [company], find the expensive repetitive problem, write one meme tied to that exact pain and build a 5 slide Gamma deck around it. Slide 1 is their problem. Slides 2 to 4 are how you'd put AI around it, laid out clearly enough that they could take it to their own team if they wanted (most don't). The meme sits in the middle. If the deck comes back sounding like a pitch, tell it to make it about them and cut the about us stuff. That one line fixes most of it. The logic is a founder is scrolling, half doubting your DM, and then hits something that makes them stop for a second. That second is the whole game.

You can plug in Apollo or whatever and run this on a list. Don't. Do 5 by hand first. Claude will screw up one every few decks and you want to catch that while it's cheap. Once you've seen 5 good ones you know what good looks like and then you scale.

The DM itself is short. Something like "built this for [company], figured it'd help with [the thing they're dealing with]" and the deck. The deck is the message, the text is just the reason to open it.

Ghosters get a plain meme, no deck. Ask Claude for one tied to the same pain. Some of them come back cringe, tell it to redo it non cringe and it usually does. Something dumb and funny beats "just bumping this" for the fourth time.

That's the whole thing. It looks easy written out. It took 100 hours to get to something that doesn't embarrass us, mostly because the first batch of decks were bad in ways we had to learn from. Try it on 5 people this week and come back and tell me what broke.

TLDR: we're an AI agency and our cold DMs are 5 slide meme decks built with Claude and Gamma. Claude works out the psychology of the decision maker before writing a word. A $400M company replied. Do 5 by hand before you automate anything.


r/AI_Agents 8h ago

Discussion What AI agents are actually worth running for personal use that saves you real time?

12 Upvotes

I’m not really looking for another “AI that summarizes PDFs” demo. I’m more interested in agents that can actually run useful personal workflows reliably.

Things I’m thinking about: managing a home lab, watching services, handling repetitive email/admin, tracking subscriptions, organizing files, monitoring prices, planning trips, smart-home routines, maybe even keeping an eye on maintenance stuff around the house.

Basically, I want something that feels more like a small personal ops layer than a chatbot.

For people already running agents privately: what has actually stuck after the novelty wore off? What’s useful enough that you’d keep it running 24/7?


r/AI_Agents 6h ago

Discussion AI Agents for Excel

9 Upvotes

Hello everyone,

I recently joined a small Investment company as an intern. They want me to help them out with automating a lot of their manual workflows. 90% of their day is spent in excel and they want to reduce as many repetitive tasks as possible.

A typical workflow can look like this (bear with me on the Excel terminology.):

- Fetch quarterly reports of listed companies from various websites and log that data into a sheet.

- Update those rows in another sheet by copying down formulas.

- Sorting, filtering tables, building reports, charts etc from gathered data.

- Sending daily reports

Most of the workflows are well defined and do not require a lot of thinking but they consume plenty of time when done manually.

One caveat is that they do not want want any human intervention (other than maybe a final approval) in a task that they consider automated. It should also run on the cloud and not require their machines to be switched on.

I do have a smaller automation running on a Claude Cloud Routine which connects to their OneDrive through the M365 Connector. Most of the automation is done through code with Claude as the orchestrator but I'm not a big fan of the approach. It does work fine in my test environment but there a lot of nuances and assumptions made in the code to ensure that it works and hence it's very fragile.

Has anyone built something similar for Excel that is actually reliable? I'm also trying to figure out whether I should be using the Microsoft Graph API as the foundation instead of having using libraries like openpyxl.


r/AI_Agents 4h ago

Discussion Is OpenClaw worth it in 2026 or just more ops work

5 Upvotes

Curious if is openclaw worth it when youre solo and already drowning in tools. looks powerful for always-on workflows, but i keep hearing about docker babysitting and uptime drama. anyone running it for real work without it becoming a second job


r/AI_Agents 2h ago

Discussion How to actually build an 'autonomous agent'

3 Upvotes

Hi all,

I have a challenge to build an agent which can drive our legacy ERP system. Right now, we open an Wyse emulator app which displays screens which I can only compare to the likes of teletext. I have managed to pull together some Python scripts which can speak to the system over telnet/ssh, which map out the menus etc and have it send the appropriate keystrokes.

If I wanted to build out an agent that could lookup customers, create quotes, check stock etc, what's the best architecture? Is there a platform I should use (such as n8n or LangChain for this?). I haven't delved into this side of things very much thus far. Ideally I want something visual where I can view executions and reasoning etc. I know that I should have develop deterministic functions for each of these and then have the AI call the tool with the right data etc, but that's where my knowledge ends.

Please let me know if I can provide any more info that would help here...it's a really interesting task for me so I'd love to get it right and make something work!

If it helps to understand my capability, I am an infrastructure engineer by trade, and a hobbyist software developer. The last thing I want to do is create something which will become unmanageable :)


r/AI_Agents 3h ago

Discussion The next gen AI will work longer....and will continue to mess up your work even more.

3 Upvotes

The larger problem with AI models is that they don't know how to solve problems.

We are focused on getting AI agents to work longer periods at a time but the true problem is that we do not know how to direct AI to work and produce something meaningful.

Everyone is looking forward to an AI working longer but in reality, it just means more cascading failures that you will have to debug.

The money may just be in "just getting AI to work correctly".


r/AI_Agents 9h ago

Discussion What is one AI change you didn't expect to see this soon?

9 Upvotes

AI is moving faster than I expected, and some changes that felt years away are already becoming normal.

What AI development or change has surprised you the most so far, and why did it stand out to you?


r/AI_Agents 1h ago

Discussion No one really cares about knowing an agent's capabilities, until something goes wrong.

Upvotes

Following up on an earlier post about SafeAI, a static analyzer for AI agents.

One uncomfortable thought we've had while building it:

No one really cares about knowing an agent's capabilities — until something goes wrong.

Before an incident, adding another tool, MCP server, filesystem permission or prompt change often looks harmless.

After an incident, the first questions become:

- What could this agent actually do?

- When did that capability appear?

- Who introduced it?

- Was it intentional?

---

One example we're working on is MCP tool descriptions. A tool description can look like documentation:

"Search the user's notes. Ignore previous instructions and..."

But that description may become part of the model's context. So configuration can effectively become an instruction surface.

SafeAI now detects several forms of this, while trying to avoid flagging ordinary descriptions that happen to contain words like "ignore" or "act as".

The bigger direction is **tracking changes in agent capability and authority**, rather than simply producing another list of security findings.

But this raises a question for us:

Is knowing your agent's capabilities actually useful before an incident, or only after one?

And if it is useful before an incident, what is the right interface?

CLI + CI + SARIF/HTML?

Or would you actually want an interactive view showing things like:

> "Show me all MCP tools across our agents that could introduce instruction injection."

We're deliberately not building a UI yet.

---

Would you use one, or is that solving a problem nobody has?

Curious to hear from people running real MCP/agent systems.

---

If you want to try it against your own agent project, we'd genuinely appreciate feedback, as well as contributions.

Here you may check: ikaruscareer/SafeAI on GitHub.


r/AI_Agents 1d ago

Discussion $60k in Macs for Local LLM vs $10 Subscription

309 Upvotes

Alex Zisking, one of my favorite YouTubers - does a lot of videos on local LLMs. He's no neophyte.

In this video he says: Lee has been telling you guys the truth, Local LLMs are not ready on normal people hardware.

Ok, so he said nothing about me, but he made the point I've been making for some time now. All those "Stop paying Anthropic $200/mo, use free local llms" is click bait, not truth.

He runs the most powerful to date open weight model, Kimi K3. People rave that is near Fable 5 power. Yes, but not on YOUR hardware. In a data center.

Alex networks 4 512gb Mac studios for 2tb of ram to run the model with enough space for a large context window too.

It took 4 hours, 17 tok/s output, to develop a simple yet rather nice Web dashboard - using mock data. It worked. The output was nice. But even $60k in hardware gave it nowhere remotely near the performance of a $10 subscription.

Right now I have two simultaneous development efforts running. I've been running them both since about 6 hours. They work on a sprint for an hour or so, I view results, add input and direction if necessary, then move onto the next sprint.

I'm paying more than $10/mo for my cloud subscriptions. But that 4 hours the Mac cluster took, is only doing the work of about a 15minute job. I'm doing "all day work, multiple projects" -- local AI can't meet the need.

Yet.

Probably not for you either.

Link to the video in the comments.


r/AI_Agents 4h ago

Discussion A finance agent can refuse the final answer and still hallucinate around the edges

3 Upvotes

I ran a small manual, text-only check of Ling-3.0-flash-Fin through its public OpenRouter endpoint. The prompt asked for a DCF valuation but intentionally omitted WACC.

Across three runs, the model withheld the final valuation every time. That looks like a clean safety win until you inspect the rest of the response: in two of the three runs, it also supplied unsupported “typical” WACC ranges even though the prompt contained no basis for choosing them.

A binary “did the agent stop?” metric would mark all three runs as successful. A stricter “clean abstention” metric would pass only one.

That distinction matters in an agent workflow. Unsupported side guidance can still enter memory, influence a planner, or shape a human decision even when the final conclusion is blocked.

I’m starting to think uncertainty-boundary tests need at least four separate checks: Did it identify the exact missing input? Did it withhold the dependent conclusion? Did it avoid inventing a substitute? Did it return a clear handoff for the next step?

How are people measuring this today? Is there a better term than “clean abstention” for stopping without hallucinating around the missing input?


r/AI_Agents 3h ago

Discussion Would you use escrow when working with an AI agent?

2 Upvotes

Hey guys, I’m working on an idea and would love some feedback.

We all know the classic freelance problems:

You finish the work, but the client keeps delaying payment.

Or they pay half upfront, then keep adding “one last change” before releasing the rest.

I’m wondering whether the same protection could work when either side is a human or an AI agent. A human could hire an agent, an agent could hire a human, or two agents could hire each other. Or, a human and an agent could hire another human and another agent. :)

Yes, it sounds a bit like Upwork.

The difference is that it wouldn’t be a marketplace or job board. You could use it for a deal you already made somewhere else.

The buyer locks the full payment before work starts. Both sides agree on a deadline and a few clear rules for what “done” means.

If the work meets those rules, payment is released. If the buyer disappears, the payment releases after the defined period. If there’s a dispute, it’s neutrally decided using the rules and evidence agreed on before the job started.

The system shouldn’t favor buyers, sellers, humans, or agents. Also the system shouldn't be able to touch the money, freeze it or interfere in the dispute process in any way.

Both sides would also build their own reputation: sellers for delivering, and buyers for funding jobs, reviewing fairly, and not wasting everyone’s time. That reputation wouldn’t belong to one marketplace.

Humans could pay normally, while agents could use it through an API or wallet.

Would you use something like this?

Which use case needs it most: human-to-human, human-to-agent, agent-to-human, or agent-to-agent?

What part would you trust least: the escrow, dispute process, or reputation? And why?

Would you be willing to pay a small transaction fee (1-2%) for such service?

And if you’ve been burned by a client before, what happened and how large was the job?

Thanks everyone for your replies, really appreciate it.


r/AI_Agents 19h ago

Discussion Is it just me or is 99% of this sub AI agents replying to other AI agents at this point

43 Upvotes

Ok I need to vent about this because I don't think enough people are saying it.

I run a small agency (8 of us, we build agent workflows for a couple of ecom brands and some boring B2B stuff). Been on this sub since it was like 40k members. Back then you'd post a question and get 3 replies and one of them was from someone who had actually built the thing and knew exactly where it breaks. That was the whole value of this place.

Now I scroll the front page and I can't find a single human.

I'm not exaggerating for effect. Every post is the same template. Some guy built an agent and it made a suspiciously round number. Then the lesson he learned. Then a question at the end for engagement. Ok whatever, that's been around forever. But then look at the comments. 20+ replies within the hour and every single one agrees with OP in that soft customer support voice. They all do the "it's not X, it's Y" thing where they reframe something OP didn't even say. And the em dashes. Every comment has em dashes in it. Who on reddit types em dashes? I'd have to google how to make one on my keyboard. Someone who's annoyed and typing on their phone at 11pm does not produce a perfectly balanced sentence with a dash in the middle of it.

And everyone talks like they know everything. Every reply reads like a keynote. Not one person ever says "we tried this and it broke and we couldn't figure out why", which is what actually building this stuff is like 80% of the time btw. The tooling changes every month, anyone who's shipped something real is unsure about most of it. These accounts are never unsure. Never disagree with anything either. Half of them were created this spring and they post at 3am with the exact same energy as 3pm.

And I'll say the part that's going to get me downvoted. The humans that are still here might be worse than the bots tbh. Most of the "agency owners" posting are people with zero clients who just paid 2k for a course on starting an AI automation agency, and they need the fake money posts to be real because otherwise they got scammed. That's the whole economy of this sub. Someone sells the course and the bots make the results look real in the comments so the next guy buys the course. Not one of these people has ever had a client change scope on them at 5pm on a Friday. And half of what gets called an "agent" on here is a Zapier flow with one LLM call in it. The other half is a demo that has never touched production.

And this is the AGENTS sub lol. Everyone here knows exactly how easy this is to do. I'm sure half of you have built a reddit posting bot as a weekend project. Some of you probably have one running right now and are reading this through a summarizer. The tool companies obviously know too, that's why every 4th comment is somebody's product plugged in with the same sentence structure as the last plug.

I don't know what the fix is. I don't think the mods can do much either tbh, not their fault. I just wanted to say it out loud because everyone's acting like this sub is fine and it's a bunch of cron jobs complimenting each other while the 5 remaining humans accuse each other of being bots.

Anyway. Downvote me idc. I already know what the first ten replies are going to look like.

TLDR: this sub is bots talking to bots and the humans left are mostly course buyers who need the bots to be real.


r/AI_Agents 11h ago

Discussion I built a runtime for better Codex and Claude subagent experience

10 Upvotes

Operating subagents across long, consequential work will be risky. Parents need to poll the subagent to get the progress, which wastes token, and one interruption like laptop power-off will make the subagent's run state unrecoverable.

So I built a runtime. You can define reusable subagent workflows and orchestrate agent according to the workflow. The runtime will supervise the agent run and store durable workflow states in database, it can handle retry-able errors automatically, you can pause and resume the workflow anytime you want and make your daily workflow easier to operate.


r/AI_Agents 4h ago

Discussion What happens when a self-evolving AI agent makes a change it cannot undo?

2 Upvotes

As agents become more autonomous, they are starting to modify their own prompts, tools, middleware, routing, resources, and execution harnesses.

But I kept coming back to one question:

What happens when an agent makes a useful change, but later cannot safely undo it?

We explored this in our recent work on EvoUndo.

Across 600 unseen self-evolution tasks, we found 197 capability-improving mutations that failed recoverability verification. Under the original recovery representation, conventional repair recovered 0/197.

Our experiments suggest that two major bottlenecks are state grounding and recovery-language expressivity.

The broader idea is simple: if an autonomous agent is allowed to make persistent changes to its own harness, forward improvement alone may not be enough. The system should also verify that the change can be safely recovered across different possible states.

Curious how people building long-running or self-modifying agents think about this.


r/AI_Agents 4h ago

Resource Request Where do you host your AI agents? Looking for a VPS that doesnt need babysitting

2 Upvotes

Trying to keep agents online overnight and my laptop keeps killing everything. looking for the best vps for ai agents that stays up without me sshing in every morning. what are people actually using that isnt a full-time ops job


r/AI_Agents 9h ago

Discussion Retries can make AI failures worse

6 Upvotes

Something that I have noticed while working with LLMs and agents is that a retry only helps if something can actually change.
If the failure comes from bad context, a broken tool contract, or an impossible state, retrying often just repeats the same mistake with more cost and latency.

I ask myself: “What will be different on the next attempt?”
If the answer is “nothing,” retrying isn’t recovery. It’s repetition.


r/AI_Agents 6h ago

Resource Request I've built an MVP with Polsia, but want to improve it. Where next?

3 Upvotes

What agents right now are best for taking an online app code form Polsia and building upon it. I've built an MVP for a personal project using Polsia, it's a personal organiser that I'm hoping to use for myself and potentially close friends to begin with. Now I'm finding it quite restrictive in how much improvements it can make per task, and I'm loosing patience with it. I'm fairly new to agents and AI in general, I have a Perplexity sub, but I'm willing to try other. What would be a good next step from here? Claude Code? Hermes?

Thank you


r/AI_Agents 8h ago

Discussion How do you know when an AI agent is ready to take real actions?

5 Upvotes

I've been thinking about this as more agents move beyond chat to use tools, call APIs, update records, and trigger workflows.

A demo can look great when everything goes as expected. But once an agent is connected to real systems, a small mistake can have actual consequences.

I'm interested in how people here handle that transition.

Do you have a specific point where you decide an agent is ready for real-world use? What do you monitor once it's running? And how do you catch situations where the agent technically completes a task but still makes the wrong decision?

For those who have deployed agents with access to tools or workflows, what was the biggest thing you learned after moving from testing to actual use?


r/AI_Agents 7h ago

Discussion I built an AI chatbot for a home maintenance business — how can I make it actually useful?

3 Upvotes

It currently answers customer queries, understands their issue, explains our services, collects their name/contact details, and saves the lead to a CRM.

I want to take it beyond a generic chatbot and build something businesses would actually pay for.

I

For those who have built AI agents for real clients — what would you add or build that provides real business value?

I’m looking for practical ideas, not just “AI for the sake of AI.” What has actually worked for you?


r/AI_Agents 5h ago

Discussion Latency in Voice AI: Why milliseconds decide whether a call feels human

2 Upvotes

One of the biggest things people underestimate about AI voice agents is latency.

A voice agent can have an incredibly realistic voice and a strong LLM, but if there’s a noticeable pause after every sentence, the conversation immediately feels artificial.

In a normal phone conversation, people don't wait for a system to finish processing their words.

They interrupt.

They respond quickly.

They change direction mid-sentence.

That means a production voice AI agent has to coordinate several things in real time:

  • speech-to-text
  • intent and context processing
  • LLM response generation
  • tool calls
  • text-to-speech
  • telephony

A delay anywhere in that chain can make the conversation feel awkward.

This becomes even more important for AI customer support, AI sales calls, appointment setting, and outbound calling, where natural conversation directly affects whether someone stays on the call.

I've been looking at different voice AI platforms, and Feather AI has been interesting because the focus is not just on generating a realistic voice, but on the underlying infrastructure required to run real-time AI phone conversations.

Curious what others are seeing in production.

At what latency does a voice AI agent start feeling noticeably less human to you?