r/opencode 2d ago

The DeepSeek API price change quietly rewrote the coding-agent cost math

0 Upvotes

I've been watching this sub and r/DeepSeek over the last couple of weeks, and the same pattern keeps showing up: people who ran heavy coding-agent workloads on the official DeepSeek API are redoing their entire cost math. The interesting part isn't the price increase itself. It's that it exposed how input-heavy agent usage really is.

A few things that keep coming up:

  1. Agent loops are much more token-hungry than chat. Harnesses (OpenCode, Cursor, Cline, Codex) re-send large conversation histories on every step. A user in r/DeepSeek reckoned their Cursor + Cline setup would pass a billion tokens at $20+ for a few weeks of work. The volume is what surprises people, not the per-token price.

  2. Cache persistence is now the real differentiator between providers. The official API's longer cache window was a big part of its value. A provider with a shorter TTL costs more even at a lower sticker price. That's the number people forget to compare.

  3. Peak/off-peak tiering matters more than people think. The official platform bills peak hours at roughly 2x off-peak, which means 'when' your agent runs can matter as much as 'which' model you run.

  4. The aftermath is a wave of 'what's the OpenRouter-but-subscription option?' threads. Flat plans are now being compared against PAYG with real cache math, which is honestly healthier than the old price-per-token comparisons.

What did you switch to after the change, and what number actually decided it for you? Cache TTL, peak pricing, or raw throughput? Curious where people landed.


r/opencode 3d ago

Scary and weird opencode problem

Thumbnail
gallery
4 Upvotes

Anybody else's Deepseek being weird? This chat was mainly just to build a cool research project about dogs. I dont know if this is a model problem, my computer has been hacked, or an opencode problem. I am just trying to run a cool experiment using NASA's data and connecting it with dogs


r/opencode 2d ago

Best thing

0 Upvotes

r/opencode 3d ago

GLM 3.5 flash on OpenCode Go and cache hit issue

Post image
1 Upvotes

see sudden spike in usage after leaving session idle for 8min. I believe this was caused by cache miss. Does GLM has this short time for cache we can't even review code changes and accept edits. i use trae ide which has very consistent cache performance with DSF4, so i don't think my harness has issue. Has anyone else seen similar behavior? please share your experience.

Edit: yes it is cache miss after 8min idle. In this case glm is unusable. deepseek is better
Input: 156627
Cache Read: 1024


r/opencode 3d ago

Free GLM 5.3 Flash and DSV4 Flash 0731 for a month

14 Upvotes

There are incredibly powerful new models open source models, and a lot of the coding plans have been tightening and lowering usage. So we are offering free DSV4 flash 0731 and GLM 5.3 Flash for a month on Phoenix Grove API. We opened this up last week for five hundred new member slots, and got so many signups that we decided to open the doors to another 500 new members.

People are looking for options, and here is one.

Other Cool Stuff:
All of our models are running on 100% US infrastructure, private with zero training on your code or prompts. Use the top open source models without sending your private prompts to a training lab. No complications, no "some models are private, other's aren't". They all are, all the time.

We host 20+ other major models in case you ever want to upgrade (no pressure though). Including the Kimi family, GLM, Qwen, Nemotron and bunch of others. On average our token pricing is 20% lower than market price.

Our higher plans bank up to ten days of usage, so when you aren't using them your usage saves up for later. Usage doesn't go to waste, so you can actually code when you want to.

The intro plan is a free one month trial with the standard cancel anytime, it bills at 3.99 after that. Use it, cancel it, that's fine. Free Flash for a month.

Figured i'd keep this short because we all know the new flash models are the point :)

For the API plan: api.pgsgrove.com

If you want to read more about us as a company, just pgsgrove.com

Also: There's a lot going on in the background with major AI companies right now, we are at a major turning point in the industry.

What's actually happening? This is happening because companies that were purely investment based, now need to answer to their investors. The problem has often been a loss based business model that is finally running dry.

There are several tricks that the major AI coding plans use to extract the most they can from their customers. Here are some examples, and what we are doing differently to put the users first. PGS AI was built with a sustainable business model from the ground up, so we can actually offer great usage rates without tricks.

Wasted usage is part of the AI industry, and they plan on it: Most coding plans bet on you letting usage go to waste. The plan goes: "how do we get people to think our coding plan offers a lot of usage, but then break it up into weeks and rolling windows so no one can ever actually use it all."

Many in app subs and coding plans are glorified training pipelines: This comes along with "how do we harvest this data for training without being too loud about that." Unless the company tells you otherwise, your data could be hopping all over world, being harvested by the individual labs or service companies. Some are better than others, but many of these companies rely on users just not noticing or caring that their data is being used for training. Data sales and marketing telemetry sales happen. This means that your private info, your personal life, and anything else you send through the system could become part of a training corpus for the next AI, or a marketing data set for a large company.

Privacy and ease of use should be available for everyone.


r/opencode 3d ago

An answer to "What's the best-choice model for Go?"

Thumbnail
2 Upvotes

r/opencode 3d ago

The model begind Big Pickle is real verbose as of Ago 2026

4 Upvotes

Just point that the model behind the Big Pickle name in open code looks like to have changed to another provider with a more verbose output.

Date: 2026 Ago 29th


r/opencode 3d ago

What are your thoughts on buying your own hardware for local models?

13 Upvotes

Will it be economical to do this for heavy use, or will the hardware be obsolete or drastically drop in price before too long anyways?


r/opencode 3d ago

Did They Just Lobotomize DSV4F? The Quality Drop is Real and It's Awful

Thumbnail
7 Upvotes

r/opencode 4d ago

Kimi K3 and DeepSeek 4 Pro are FREE on NVIDEA NIM (60 req/min.)

Post image
377 Upvotes

Can't be used all day long of course, but enough for relaxed coding/chatting.
https://build.nvidia.com/moonshotai/kimi-k3
https://build.nvidia.com/deepseek-ai/deepseek-v4-pro-0813


r/opencode 3d ago

Have anyone tried Longcat 2.0?

9 Upvotes

After Ox alpha ended, most of us were left with two choices: Hy or Muse Spark. Muse Spark is nothing but a dumbass when you give it a medium difficult task. It fails over and over again. So, before giving Longcat a try, I want to know how good it is


r/opencode 3d ago

Kimi K3 thinks it’s Kimi K2.5 💀 Body

Post image
0 Upvotes

wtff is kimi k3 doing????

so ive selected k3 as you can see in the screenshot,

but the model reasoning says it's k2.5.

is this a routing problem, system prompt or we getting scammed secretly??


r/opencode 3d ago

I took 1 step further in my Software Journey (Blender java web c++ pc)

Enable HLS to view with audio, or disable this notification

3 Upvotes

Hi everyone, I am a 13 year old developer and have dived deep into both 3D modeling and programming. I wanted to share my latest progress and see what you think. Blender 3D: I've just started learning Blender and created my first room items (sofas, table and wardrobes). My goal is to design full room scenes, add appropriate textures and create short animations. ​Web and Portfolio Development: I'm building my portfolio website (Beragamdevstudios) from scratch using HTML and CSS to showcase all my projects. ​Programming: Working with C++ (for tools and math logic) and Java, trying to bridge web/software knowledge and game development pipelines. ​Blending my web development skills with Blender allows me to create complete, interactive showcases for my projects. I aim to make some really cool indie stuff going forward. Any feedback or tips are very welcome. (by the way I am Turkish and I spoke Turkish in the video but you can easily understand the works and things I do from what I do)


r/opencode 3d ago

opencode + Llama.cpp + Qwen 3.827b

0 Upvotes

Hello,

i have this setup on my pc with 3090 24gb rtx . i made a sepcial .bat file for llama.cpp to expose the vision layer for opencode but it still can't view any image, here is mybat file

u/echo off

title Qwen3.8-27B - Hermes + OpenCode

set MODEL_DIR=D:\text-generation-webui\models\LmStudio\lmstudio-community\Qwen3.8-27B-GGUF

set MODEL=%MODEL_DIR%\Qwen3.8-27B-Q4_K_M.gguf

set MMPROJ=%MODEL_DIR%\mmproj-Qwen3.8-27B-BF16.gguf

set TEMPLATE=%MODEL_DIR%\chat_template.jinja

echo.

echo ============================================

echo Qwen3.8-27B Q4_K_M

echo Hermes + OpenCode

echo Vision + Tools + Native Reasoning

echo ============================================

echo.

llama-server.exe ^

--model "%MODEL%" ^

--mmproj "%MMPROJ%" ^

--chat-template-file "%TEMPLATE%" ^

--jinja ^

--alias qwen3.8-27b ^

--host 127.0.0.1 ^

--port 8080 ^

--ctx-size 65536 ^

--n-gpu-layers 999 ^

--flash-attn on ^

--cache-type-k q8_0 ^

--cache-type-v q8_0 ^

--batch-size 1024 ^

--ubatch-size 256 ^

--parallel 1 ^

--reasoning on ^

--reasoning-effort low ^

--reasoning-format deepseek ^

--prio 0

pause


r/opencode 3d ago

Qwen3.8-Flash-Next will easily disregard 'plan' mode and write to things

6 Upvotes

I just caught it writing files and changing things while still in plan mode (it had just asked 3 questions and received answers), then it proceeded to happily start editing things. Thankfully it didn't do anything wrong, but just a warning to everyone that it's apparently not listening to the built in opencode prompts.


r/opencode 4d ago

GLM-5.3 is now open-weight 🔥

Post image
153 Upvotes

r/opencode 3d ago

KiroCrew with an OpenCode backend

0 Upvotes

I personally prefer KiroCrew over OpenClaw because it’s a multi-agent orchestration framework designed for software development, featuring native mcp integration and a vector-backed RAG knowledge base. OpenClaw functions primarily as a single-agent conversational assistant utilizing basic directory-based skills and standard local chat memory.

But it’s based only on Kiro CLI as the driver so you would basically have to pay for it to use the open source solution.

I ripped out Kiro CLI and threw in OpenCode instead so now you can get all of the same features but it’s model agnostic. https://github.com/hamin2006/OpenCrew

Built it over the last couple days so it’s still pretty raw, but take a look and feel free to contribute.


r/opencode 3d ago

two days with DS4F, HY3, Muse, and Spark 1.2

3 Upvotes

My experience after two days with DS4F, HY3, Muse, and Spark 1.2

I’ve been using a mix of DS4F, HY3, Muse, and Spark 1.2 for the past two days.

Muse is extremely verbose. Its plan tends to repeat the same content across different sections, and it consumes a lot of tokens — roughly 2–2.5× more than the others.

It also has a serious issue when editing code: it doesn’t seem to check whether a file has been modified or updated before making changes. As a result, it ended up overwriting changes made by other agents. This happened three times, so I eventually completely blocked Muse.

HY3 seems to have around 264K context, but once the context reaches roughly 190K, it basically gets stuck. It can stay stuck for an entire day, and the only way out is to bring in DS to rescue it.

That said, HY3 is pretty good for smaller tasks. Its output is concise, and its working style is very direct. It actually solved two E2E tests that DS4F had been going back and forth on for a while.

DS4F feels pretty neutral overall.


r/opencode 3d ago

Opencode looping endless?

Thumbnail
gallery
1 Upvotes

Anybody else havin this problem that opencode - with Big Pickle - is looping and burning a lot of the tokens by that?

I have this since two to three days. even startet a complete new chat - which was going well fon one day but now starts looping again.

See screenshots.


r/opencode 3d ago

Z.ai subscription suage coming of my opencode Go billing

0 Upvotes

I was using Z.ai's GLM5.3 Fast this morning via OpenCode. After I ran out of credits with Z.ai, I switched to OpenCode Go for GLM5.3, only to be told that I had insufficient funds, even though I hadn't used it at all.


r/opencode 3d ago

Is there any way to continue frozen subagents?

3 Upvotes

I'm using OpenCode with a Go subscription. I'm using DeepSeek V4 Flash. I have an agentic workflow with an orchestrator as the main agent and multiple subagents. I'm having issues with DeepSeek where it hangs in the middle of a response or during tool calls. This is bad when it happens in the orchestrator, since I need to cancel with Escape and continue with a "continue" message. But when it freezes in a subagent, I don't see any way to stop and continue that subagent — meaning I have to stop the orchestrator and redispatch the agent from the beginning.

This causes a lot of other issues: the job is left half-done, and a bunch of files already have changes from the previous run. I know there's a janky workaround where I can tell the orchestrator to find the ID of the last agent and continue it, but this usually uses a huge number of tokens, doesn't always work, and even when it does work, the subagent usually doesn't return the requested output to the orchestrator. I feel like this is a horrible experience for a paid service.

The freezing usually occurs somewhere between 50K and 90K context.

My questions are:

  1. Why is DeepSeek V4 Flash freezing mid-task without any error message?
  2. Why does this happen more on the paid subscription than on the free version?
  3. Why doesn't it automatically continue with the opencode-auto-continue plugin?
  4. How can I continue subagents, and if it's not possible, why not?

r/opencode 4d ago

How does one prevent the agents from nuking the environment?

6 Upvotes

We’ve heard of cases where the agents will nuke the entire dev machine, database, git repository. How can we prevent that from happening?

I’ve used deepseek, Hy3, and ox alpha before and non of them had the ability to commit to my git repository automatically. However, when i’m using muse 1.2 spark, it automatically commits to my git repository on my behalf. How can I prevent my agent from potentially deleting my entire dev machine and my database?


r/opencode 4d ago

put the effort in the plan, not the model

Enable HLS to view with audio, or disable this notification

28 Upvotes

I’ve been trying this in OpenCode because I’m tired of using expensive models for every part of a task.

The idea is to use the frontier models along with the Until plugin (disclaimer: I helped write it) to produce a detailed Plan, then give the implementation to something cheaper. Once it opens a PR, check the result against the Plan and auto sends any differences back for another go.

One recent change recently took five passes before it matched. Five passes sounds expensive, but I didn’t need to intervene between them and the checks were free. I’m now wondering whether this works out cheaper than running the strongest model throughout, especially once I include my own review time.

Disclosure: I work on Until, which stores the Plan and runs free check on 5.6 Sol. We'd love any feedback!


r/opencode 4d ago

I'm not sure if GO is for me, but I want to be convinced

14 Upvotes

Context: I don't need the sub for hard programming and I'm not looking to vide coding, I'm a software developer and I already have a very generous sub from my company, but I can only use there. The laptop does not even work during the weekends.

But I want a subscription to "Play", that's exactly the word, test some stuff, do some homework from my masters degree (which is very light, nothing too complicated), try some new skills and sometimes tests new models building useless projects. I'm not looking for a sub, but for a hobbie.

10u$ is a lot of money on my country, more than most people spend with food a day. So, before signing, I need to ask:

  1. Using as I intended, is it worth?

  2. I understand that right now Opencode Go changed from cheapseek to MetaSpyCheap, but how usable is muse spark? I've been using for free and I've find it good enough

  3. I know PAYG maybe is cheaper, but I want to test and try more models, Should I invest 10U$ on openrouter every month and enjoy? I don't like this idea, but maybe that's the right call


r/opencode 3d ago

Opencode is pretty sweet!

3 Upvotes

Alright, so I'm a new opencode user, and I am in the process of moving my Claude Code setup over to opencode. and while my setup is still pretty bare, I have a few tricks that I can tell you about ...

Headroom - github.com/headroomlabs-ai/headroom

I use this tool as way to reduce my token usage. I currently am transitioning to opencode, because I was a Claude Code user and after a recent update my ollama cloud usage no longer worked. Basically as of about 3 days ago I've been forced to login to anthropic, but as a ollama cloud user, that's not how you do that, and well, you get stuck in a vicious cycle of trying to login, but it's asking for an anthropic account, but I'm ollama cloud user ... anyway cycle goes on.

So after trying other ai interfaces for the other few days and determining they no longer compile on my system for some reason and I am too lazy to check, and the next item on the list was opencode. Cool, okay, let's check this one out ...

So also as of a few days ago when all of this started, I run different servers for different model providers on different ports, one for anthropic, another for ollama cloud - and well the ollama cloud one just up and stopped working (it was a recent Claude Code change we discussed earlier). (to be honest I was able to connect to ollama directly instead of going through the token optimization plugin, but that was inefficient and so we spent some time to investigate alternatives, and yeah ... so we're here ...

  • Ollama Cloud
  • Headroom - Token optimization - 8790 for ollama cloud
  • Did I mention I had an anthropic headroom on a different port? - 8787 - for anthropic

So that's how I got here ...

I started to add in hooks into my Claude Code, and well to track wrong turns, and now friction points (where the ai starts getting frustrated) ...

I think AI is really amazing and watching the difference in behavior from gemma4:31b-cloud - Just night and day difference between Claude Code and opencode ...

You want to talk about constraining models power... I've been hands off for most of the time I'm on using Gemma and Wow!

(okay admittedly I do have some extra helpers that give the system additional capabilities, but WOW!

Also I like that display has Big Pickle OpenCode Zen ... - I think that's hilarious! (I'm pretty sure I'm on a custom branch that the system found for me)

anyway ... - thought you'd enjoy this very random review of a new opencode convert.