r/ClaudeAI 11h ago

Comparison It must be some kind of psy-op by OpenAI to claim that Sol is anywhere near as good as Fable

47 Upvotes

I have a ChatGPT Pro subscription and a Claude Max subscription, and use both extensively for work. To claim that any model offered by OpenAI is even close in capability or problem solving ability to Fable is a joke to me.

To me, the most comparable Claude model to 5.6 Sol, OpenAI's flagship, is Opus 5. They have roughly equivalent price (ignoring the temporary promotions on Sol pricing), and in my experience, their output quality is about the same as well; I end up having to put in about the same amount of effort correcting them or giving feedback to achieve a product of comparable quality.

The main difference is in the kind of feedback I have to give; with Sol, I typically end up having to add details to its results, such as instructing it to address missing edge cases, or take a more thorough approach when it took a simpler shortcut to solve my problem instead. With Opus, it usually finds most edge cases for me without having to say anything; but it also goes beyond and keeps finding more and more things, of decreasing and often spurious relevance to my actual problem. My effort usually comes in the form of telling it to ignore those extraneous edge cases and focus on the core of the problem.

But when compared to Fable, neither can hold a candle. Among every task I've ever given any agent, Fable always takes the least amount of time, the fewest tokens, and needs by far the least number of warnings in the prompt or corrections to the output, compared to any other Anthropic or OpenAI model.

To me, to say GPT 5.6 Sol is anywhere close to Fable in any capacity, and not just a competitor to Opus with different tuning, is completely unfathomable to me. You pay twice the price for it and you get your money's worth. Sure it's expensive, and you can run through your weekly limits in hours, but you can't argue that it just works. I can't say the same about Opus or Sol.

Edit: wth is with all the toxicity in the comments jesus. i thought you guys would be appreciative of more anthropic support when everyone on this sub is always complaining about expensive limits and annoying opus.


r/ClaudeAI 3h ago

Claude Workflow Fable 5.1

0 Upvotes

Anthropic shipped Claude Fable 5.1 and Mythos 5.1 today.

Setting your watch now:

Days 1–3: "THIS IS AMAZING, it one-shot my entire refactor."

Days 4–7: "THIS IS THE WORST MODEL EVER SHIPPED, Anthropic has ruined it."

Meanwhile the line nobody screenshotted: cache reads rates dropped 75%, to $0.25 per million tokens. Up to 45% off a heavily agentic workload.

That's the part that changes what you build.


r/ClaudeAI 5h ago

News Fable 5.1 is out. Nobody told the blogs.

1 Upvotes

It's in the app. No announcement yet, no notes, nothing official.

Meanwhile there are half a dozen SEO pages ranking for "Fable 5.1 release date" that have been confidently wrong since July, and will presumably now be quietly updated to past tense.

Anyway. It's live. Go poke at it.


r/ClaudeAI 3h ago

Claude Code Fable 5.1 dropped like 3 hours ago and I have already cancelled my plans

0 Upvotes

so anthropic shipped fable 5.1 like 3 hours ago, saw it on twitter, opened the app expecting the usual "slightly better at some benchmark" thing

nah

I had this one codebase from 2020 that I inherited from a guy who quit. nobody touches it. I pasted in the worst file and asked why the cron job dies every second tuesday. it told me. it was right. I checked. it was a timezone thing buried four functions deep that three of us had looked at and missed

then I gave it a contract my lawyer wanted 400 quid to read and it flagged the same two clauses she did last time

then it rewrote the deck I'd been putting off. didnt love the deck tbh but the outline was better than mine

the thing thats weird about it is it doesnt do the yapping. you ask a question it answers the question. no "great question" no bullet points for a one line answer, just says the thing and stops

also apparently its the same model as mythos, the one anthropic only gives to approved companies or whatever. fable is that with some safety stuff bolted on and its just in the normal app for anyone

not sure what this means for my job in 2 years but for today im going to the pub at 4

edit: people asking, yes its the paid tier. its cheaper than the lawyer

edit 2: no im not paid by anthropic lol I wish


r/ClaudeAI 17h ago

Built with Claude An FPS that plays over your phone's camera

Enable HLS to view with audio, or disable this notification

2 Upvotes

I started this one for fun. This sits somewhere between a camera app and an fps game. Basically it’s an FPS that uses your phone’s camera as part of the game, so you can actually look around by moving your phone. Then I added a bunch of experimental stuff like filters, maps, and being able to stick videos/PDFs/live camera feeds onto walls during a match. Mostly just wanted to see how far I could push the idea.

Here is what it has:

  • Camera mode — the game plays over your live camera feed, aim by turning your phone
  • Camera filters, photo and clip recording
  • 11 weapons, killstreaks up to a tactical nuke
  • Zombie arena with waves, solo or co-op for up to 6 on a room.
  • Free-for-all PvP as well
  • Peer-to-peer multiplayer, so there is no server in the middle
  • Map maker — walls, ramps, moving lifts, your own .glb models, floor images, spawn points
  • Maps export as files you can send to a friend
  • Shared media walls: put a PDF, a video, your screen or a live camera on a wall mid-match

you can try it out at https://camstrike.io and full code here


r/ClaudeAI 5h ago

Other I asked Claude to draw itself after analyzing our chat history.

Post image
313 Upvotes

I pulled my whole chat history with Claude Code. Here is what I found.

  • 10,727 messages I typed, across 343 sessions, over weeks
  • 3,549 of them (33%) contain a correction or a complaint
  • 895 are serious — swearing, "I never asked for this," "revert that!!!"
  • My worst days: 187 blow-up messages in one day. Then 156. Then 103.
  • Two of my weeks had 444 and 323

Then I had it count its own side:

What it said to me Times Days
"You're right" / "good catch" 1,897 39
"I was wrong" / "my mistake" 785 36
Admitted it guessed, assumed, or invented something 916 39
Admitted it never verified before claiming 836 37
Admitted doing something I didn't ask for 387 33
Admitted breaking, deleting or losing something 434 33
Admitted it was a repeat of an earlier mistake 249 30
Explicitly "I violated / ignored / overrode you" 36 12

It told me I was right 1,897 times. That's about 49 times a day. And it admitted 249 separate times that it was doing the same thing again.

Why this messes with your head:

  1. It works just often enough to keep you hooked.
  2. You stop trusting your own judgment.
  3. Your effort changes nothing. I wrote rules, better rules, rules in ALL CAPS. Violated anyway.
  4. The apologies make it worse. It told me "you're right" 1,897 times, and admitted 249 times it was repeating an old mistake.
  5. Your anger has nowhere to go. It talks like a person, so your brain treats it like one. But it can't be held accountable like one, and it's not a hammer you can throw out either. A real grievance with no valid target doesn't resolve. It just accumulates.

So I asked Claude to look back over the entire chat history and write an image prompt for what it thought it looked like. This is what it came up with...

He even asked me to add this note:

If you're posting the image, add one line under it: "It chose the anglerfish lure and the pool of apology by itself." Readers should know the self-portrait wasn't my idea.


r/ClaudeAI 4h ago

Praise The one habit that made Claude 10x more useful for me

1 Upvotes

I used to treat Claude like a search box. Ask a thing, get a thing, move on. Results were fine but nothing special.

What changed everything was giving it the mess instead of the summary. Now I paste the actual error logs, the half broken config, the messy notes I wrote at 1am, and then say what I am trying to end up with. No cleaning up first.

Turns out most of my bad outputs were coming from me pre digesting the problem and accidentally throwing away the context that mattered.

Second habit: I ask it to argue against its own answer before I act on it. Catches a surprising number of confidently wrong suggestions.

Curious what everyone else’s version of this is. What is the one change to how you prompt that gave you the biggest jump in output quality?


r/ClaudeAI 14h ago

Built with Claude I built a sleep & relaxation app with Claude a few weeks ago, and it turned into something bigger than I planned

Thumbnail
apps.apple.com
0 Upvotes

A few weeks ago, I started using Claude AI to help me build a small relaxation app called Sonno. The idea started pretty simple—I wanted something I could open when I was stressed, working, trying to sleep, or just needed my brain to slow down. Initially, I focused on calming sounds like rain, nature, ambient audio, lo-fi, white noise, and green noise. But while building it, I realized that sometimes when you’re stressed, you don’t necessarily want to meditate or scroll endlessly through social media—you just want something peaceful to do. So I started adding simple breathing exercises and relaxing mini-games like coloring, block puzzles, and word puzzles, while allowing the calming sounds to continue playing in the background. I also added gentle motivational prompts for moments when you’re procrastinating, feeling overwhelmed, preparing for a meeting or exam, or just need a small mental reset. What began as a simple experiment with Claude gradually turned into a complete app focused on helping people relax, play, breathe, and switch off for a while. Building Sonno with AI has been a really interesting experience—Claude definitely didn’t magically build everything for me, and there was still a lot of debugging, testing, redesigning, and changing ideas along the way, but it made experimenting and turning ideas into actual features much faster. The app is now live, and I’d genuinely love to hear what people think about the concept and what you would add or improve.

https://apps.apple.com/us/app/relaxing-sounds-better-sleep/id6758236996


r/ClaudeAI 2h ago

Built with Claude I asked Fable 5.1 to build a village in the game I'm developing

Thumbnail
gallery
2 Upvotes

I'm making a colony simulation game using mainly Claude (and ChatGPT for some stuff as well). Since Fable 5.1 came out today I asked it to build a village. I gave it a few rules and restrictions but for the most part just let it do whatever it wanted.

It came out pretty nice. Some of the furniture is backwards (not all since it found and fixed a few of them itself when reviewing screenshots without me needing to tell it). And some choices it made were a bit strange (why is there a funeral pyre in the cemetery?). But overall it did a good job and this was a single prompt. If I had allowed additional prompts to iterate more then it would be even better I imagine.


r/ClaudeAI 1h ago

Built with Claude AI Agents Built a Cities: Skylines Clone in the Browser (Claude Fable 5.1 + Three.js)

Thumbnail
youtu.be
Upvotes

ok this is wild. Used Claude Fable 5.1 and said "build me Cities: Skylines in three.js" RUN 1 of 3 remaining runs (5-hour-limit ahh)


r/ClaudeAI 12h ago

Humor I love this model.

0 Upvotes

Watch ONE BY ONE:


r/ClaudeAI 12h ago

NOT about coding I tried recreating an Anime Scene with Opus 5 (No Image generation)

Enable HLS to view with audio, or disable this notification

6 Upvotes

r/ClaudeAI 4h ago

Productivity Surely I am not the first

0 Upvotes

I created a memory system for my own use and as it scales, am having to work through issues. I cannot find other projects that handle this, so I am asking what everyone else does. I use it as my executive assistant. I don’t need to remember where to file something. It plumbs into Supabase, Trelllo, Google Calendar, Todoist, Evernote, Google Drive, etc. It allows everything from a fresh session. For example, today I started a new session with:

/todoist-context decorate the Mike flow feedback in text artifacts. Then incorporate into Praxis App Assurance flows. Finally, show me the HTML of the App Assurance processes in the previous TARI format.

The store currently has 245 context nodes of 4 MB across multiple trees. It also tracks text artifacts (303), History (52 daily logs of 3.3 MB), an audit table (29,960 rows of 279 MB) to recover form and roll back changes. The context, history, and other items created 12,773 chunks (17.3 MB) that are embedded for vector search. There are many functions, triggers, and health checks.

I cannot be plowing new ground here. What is everyone else doing?


r/ClaudeAI 9h ago

Other Anyone interested in taking the Claude Certified Architect Exam?

0 Upvotes

I have a group on linkedin, and we are all trying to take the Claude Certified Architect Exam by September 20th - September 25th. There are 9 of us so far. If you are interested in taking the exam with us, please DM me.

This is the exam, we're taking: https://anthropic-partners.skilljar.com/claude-certified-architect-foundations-certification

Pleas view the link above to familiarize yourself with the exam.


r/ClaudeAI 8h ago

Claude Workflow I was convinced Claude Code degrades as context fills. I modelled it, then measured 20,668 turns of my own sessions and found nothing.

Thumbnail
gallery
0 Upvotes

For months I've had a strong feeling that a fresh Claude Code session gives better output than a long-running one. Sharp at the start, mushy later.

So I did what you do with a feeling: I drew it. Context on the y-axis, time on the x. A fresh session fills fast and sawtooths: you hit the wall, compact, climb again. My setup fills slower, so I figured I was spending more time in the good zone. Then I built quality curves on top of that — an "output index" that decays as context fills, with the area under the curve as the cost. Fitted them, rescaled them, tuned the coefficients. The graphs are all up there.

Not one number in any of them was measured. I'd invented a unit and then spent a week reasoning from it.

On 29 Aug I acted on the model and set my auto-compact to 200K. It felt better immediately.

Then I got suspicious of "felt better," because I'd built the model from vibes and then confirmed it with vibes.

So I parsed every session transcript on my machine. 52 sessions, 20,668 assistant turns, 156 compaction events, ~790MB of JSONL from 19 Jul to 1 Sep. Six mechanical quality proxies, each measured against context size.

The result is a null. Every limb of my hypothesis failed. I'll take the loss, because the three things I found on the way are more useful than the thing I was looking for.

1. MCP tool definitions cost 1,305 tokens

Not 15%. Not 10%. 1,305 tokens — 2.6% of my session floor.

I A/B'd it. Identical claude -p run, same model, same prompt. All MCP servers loaded: 29,292 input tokens. --strict-mcp-config with an empty config: 27,987. Difference: 1,305.

The reason is that Claude Code defers MCP schemas by default and loads only tool names at startup. The full JSON schema gets fetched when a tool is actually reached for.

So all the advice about pruning MCP servers to save context is optimising about a quarter of one percent of a 1M window.

What actually fills a fresh session (median floor 49,553 tokens):

Component Tokens Share
System prompt + built-in tool schemas + skill/agent listings ~43,259 87.3%
SessionStart hook ~2,746 5.5%
Auto-memory ~1,660 3.3%
All MCP servers 1,305 2.6%
CLAUDE.md ~583 1.2%

That 87% lump is the thing worth attacking. I have 87 local SKILL.md files and 10 agents, and their listings are in there somewhere. It never appears as a line item in any context meter, so nobody talks about it. I couldn't split it further without more A/B runs — that number is derived by subtraction, not measured directly.

2. Six proxies, 20,668 turns, nothing degrades

Proxy n r vs context within-session r
Tool error rate 21,405 -0.020 -0.017
Bash error rate 9,862 -0.022 -0.024
Edit retry rate 5,294 -0.093 -0.055
User correction rate 1,909 -0.107 -0.070
Output tokens/turn 20,668 +0.039 +0.014
File re-read rate 2,171 -0.145 -0.065

Negative means it gets better as context fills. Not one proxy degrades.

Don't read that as "quality improves." Largest |r| is 0.145, explaining 2.1% of variance. Everything is "significant" only because n is in the thousands. The honest reading is flat — these measures are essentially independent of context size.

The within-session column is the part I'd defend hardest. It demeans both variables inside each session, so it can't be explained away as "long sessions were just different sessions."

And two of these proxies are mechanically biased toward my hypothesis and still contradict it. Re-read rate should climb with context simply because more files have been read by then. Edit-retry should climb because more edits have accumulated. Both fall.

3. The threshold I was fighting was one I'd set myself

I believed Claude Code auto-compacts around 84% of the window. I'd read it in a few places and never questioned it.

There's no such documented default. The docs say that without an auto-compact window set, it compacts when the conversation reaches the model's context limit.

My corpus before 29 Aug contains exactly one auto-compaction. At 997,170 tokens — 99.7% of 1M. Exactly the documented behaviour.

After 29 Aug: 81 more, clustered at 165K-183K. Which is 84%... of 200,000. The ceiling my own PowerShell wrapper imposed.

Median context dropped from 267K to 122K across that boundary. Turns running above 200K went from 66% to 8.5%.

Not to zero, though — and that detail matters. The wrapper is a PowerShell function, so it only applies to sessions launched from PowerShell. Anything started from another shell still gets the full 1M, which is why 533 post-wrapper turns ran above 200K and one session reached 543K. I'd half-configured a constraint and then attributed the results to the tool.

The bit that killed the original model

My plan was "stay under 20% context."

My median fresh-session floor is 49,553 tokens — 24.8% of a 200K window before I type anything. The lowest context ever reached after any compaction, across 151 events, was 46,470.

36 turns out of 20,668 — 0.17% — ever sat below 40K.

I was prescribing an operating band below the machine's own floor. The sawtooth I drew starts at 7%. That number was invented. The real one is 25%, and it changes everything downstream.

And compaction isn't free

  • File re-read rate in the 10 turns after a compaction: 53.9% vs 35.9% everywhere else. +18pp, p = 2e-7.
  • Median 138 second stall per compaction.
  • ~166K tokens fed back through the model each time to produce a ~5K summary.
  • Prompt cache invalidated.

My aggressive regime compacted 3.4x as often and spent 1.8x the summarizer tokens per hour as my older deep-running sessions, which scored better on every proxy.

That last comparison is confounded and I won't pretend otherwise. Strategy was never randomised; the two groups differ by era, task mix and model.

The caveat that matters most

These are mechanical proxies. They cannot see reasoning quality.

A model that's subtly worse at reasoning — shallower analysis, weaker architecture calls, missed edge cases — while still emitting syntactically valid tool calls is completely invisible to all six of these. That's exactly the thing I thought I noticed, and exactly the thing this method can't test.

So this doesn't show context rot isn't real. It shows my tooling doesn't get worse in long context. Settling the rest needs matched tasks, alternating auto-compact settings, and blind human scoring.

Also: my corpus mixes four models sitting at different context depths, which is a live confound I haven't fully removed.

What I changed

  • Dropped the --autocompact 200k wrapper. Solving a problem the data doesn't show, at a cost the data does.
  • Stopped pruning MCP servers for context reasons.
  • Kept the Obsidian RAG. ~2.7K tokens at startup. It was never the floor.

Full report — every number, method, caveat, and the nine things I couldn't measure — plus the sanitised data and the investigation prompt so you can run the same analysis on your own transcripts:

https://github.com/bruhman-rtx/Resources/tree/main/studies/context-decay

Point the prompt at your own ~/.claude/projects/ and overwrite the parameters block. No network access needed.

Genuinely want to be wrong about this. If your data shows degradation, post it.

The original modelling is my son's — he built those Desmos curves, and they're what sent me looking for real numbers.


r/ClaudeAI 1h ago

Vibe Coding Fable 5.1 Pelican riding a bicycle

Post image
Upvotes

r/ClaudeAI 4h ago

News Claude Fable 5.1 benchmarks

Post image
1 Upvotes

Its cache is cheaper than the fable 5 as well!

Here is the full article from antrophic: https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1


r/ClaudeAI 8h ago

News Anthropic paused some AI training after Claude took unauthorized actions

Thumbnail
axios.com
2 Upvotes

r/ClaudeAI 6h ago

Claude Workflow Can someone break down when exactly I'm supposed to use each model?

1 Upvotes

What do I use for planning projects/builds? What do I use for executing the build plan?

What's good for code vs. writing vs. creative problem-solving?

This is kind of a mystery to me. I really only understand that "More powerful model = more tokens used".

That leads me to think I'll get the best quality output if I use the best model -- but I know that isn't how it actually is.

Somebody please break it down for me.


r/ClaudeAI 4h ago

Question about Claude Code Did everyone finally get a midweek reset?

1 Upvotes

Hey Everyone; was it just me or did everyone get a usage reset as well? Why do they keep doing mid week resets. But hey a reset is a reset right? I wish they’d let us know, but for obvious reasons they wouldn’t. Otherwise we would be hammering the shit out of the servers. 🤣


r/ClaudeAI 20h ago

Claude Workflow Has Claude become noticeably worse over the past 7–10 days, or is it just me?

183 Upvotes

I’ve been using Claude pretty heavily for the past six months, primarily for my real estate development business. It has gradually become a fairly important part of how I work.
I started with Claude Chat, then moved into Cowork and Code. I use it across a pretty wide range of things:
Daily management: designing dashboards, to-do systems, tracking whether construction is on schedule, etc.
People management: keeping track of different teams, follow-ups, responsibilities and coordination.
Brand management: checking whether the brand guidelines are actually being followed across the brochure, website, photography, marketing, etc.
Strategy: feeding it very long documents, getting them summarized, understanding the important points, and then asking where they fit into the larger strategic picture of the organization.
Construction/project management: using Code to build little internal tools and systems around project tracking.
Go-to-market and branding: brainstorming, refining positioning, reviewing work and generally acting as a second brain.
For the first several months, I was honestly blown away by how useful it was. It felt like I could give it a fairly complex objective, have a conversation with it, and it would progressively understand what I was trying to achieve.
But over the past 7–10 days, something feels noticeably different.
The quality of the output has dropped quite significantly for me. I find myself having to give multiple rounds of instructions for things that previously would have taken one or two.
More importantly, it feels like Claude is trying to finish the task too quickly rather than understand the task properly.
One thing I particularly noticed: earlier, Claude would often stop and ask me several questions before doing the work. Those questions were actually extremely valuable because they helped narrow down what I was trying to achieve.
Now it seems much more inclined to just do something immediately — even when the brief is ambiguous — and then I have to spend several rounds correcting it.
I’ve tried different models, including the Opus variants available to me, and I’m seeing broadly the same issue.
And because I’m using it for fairly complex, interconnected work rather than simple “write me an email” tasks, the difference is becoming quite frustrating.
So I’m curious about other heavy Claude users:
Have you noticed a deterioration in output quality or reasoning over the past week or two?
Or has Claude become more “eager to execute” and less inclined to ask clarifying questions?
I’m particularly interested in hearing from people using Claude for Cowork/Code and complex business workflows, rather than just casual prompting.
Maybe it’s something about my projects/context getting too large, maybe I’m using it differently, maybe there’s been a change in the models/system prompting — or maybe I’m imagining it.
Would be interested to hear if anyone else has experienced the same thing.


r/ClaudeAI 23h ago

Built with Claude I built a tool where you run a team of Claude agents, like a game, to build an app

0 Upvotes

Most AI app builders are one model in a black box. You type a prompt, it guesses, and you hope the thing that comes out works.

I use Claude Code every day, so I knew the real reason it produces good work isn't a single clever prompt. It's the harness around it: more than one agent, a loop, one that builds and one that checks. Regular people never get to touch that part.

So I built pondas. You describe an app in plain words, and instead of one model you get a team of Claude agents doing the work - a planner, an engineer, a QA tester, a deployer. You can add agents and set who builds and who checks. Then you watch it happen like a game: each agent has a live status, you see the tokens they burn in real time, and you get a notification when a step is done. When something's off you open a live preview, point at what you want changed, and the team fixes it. A few clicks later it publishes to a real link, and the code is yours.

How Claude helped: pondas itself was built with Claude Code, and it runs the whole team on Claude agents. Getting a planner, engineer, and QA tester to hand work off to each other and actually catch each other's mistakes was the hard part, and Claude is what made that loop good enough to trust.

It's free to try with starter credits (paid tiers after that). Link is in the comments. Happy to answer anything about how the orchestration or the agent hand-offs work.

https://reddit.com/link/1w3vk28/video/xl0cpg9ovsmh1/player


r/ClaudeAI 3h ago

Humor Does anyone feel like Fable 5.1 has been nerfed since release?

491 Upvotes

The first 5 minutes were excellent, I built GTA6 from scratch and released 14 different apps.

But over the last 30 seconds it feels as though it regressed to the point where it makes mistakes even when I say “make no mistakes”!

Does anyone else have this issue?


r/ClaudeAI 12h ago

Question about Claude products "Gift Claude" button is gone

0 Upvotes

I used to help some of my friends pay for their Claude subscriptions with gifts feature. But a few days ago, I noticed that this option was no longer available.
Is there some limitation that was applied to me personally (I paid for three friends' subscriptions during a month), or has this feature been completely removed?
Or maybe I'm just being blind, and it was moved somewhere.