169
u/IAM_274 4h ago
So according to Anthropic: Opus 5 beats Fable 5 on almost everything by a noticeable margin. Hmmm. Makes me question the integrity of these benchmarks in the first place.
53
u/Formally-Fresh 4h ago
Yeah that’s weird to me personally fable feels about 5x better
18
u/ThatOtherOneReddit 4h ago
My experience is they are similarly capable but OPUS seems to wildly degrade even over medium length sessions so fable is much more reliable doing anything significant.
4
u/crimsonpowder 2h ago
Looked into Opus degradation, and grounded it in real signals instead of hand-waving. Long context can in practice be a contributory directional signal to context coherence, but there are other CDS I'm ready to look into. Todos updated. Your call.
28
u/Hediak-Chigashi 4h ago
Also Sol is apparently lower than Opus 5 in everything. The OPUS 5!!! 😂😂😂
16
u/PlasmaChroma 3h ago
Those Sol numbers are highly sus. The agentic coding one at the bottom at least seems more believable.
6
u/mathtractor 3h ago
Its Mythos 5 feeding Opus 5 RL agentic work on tasks associated with the benchmarks. Fable is a clear superior intelligence but opus is dogged and persistent and RLed hard as F.
5
u/Vivid-Snow-2089 2h ago
this alone makes the entire chart pretty much magic smoke garbage, cause opus 5 sucks compared to fable 5
•
•
u/Key_Reading_9664 33m ago
makes me question veracity of the feedback on reddit and social media. Most of the complaints I saw (and had myself) were around its communication style, not the results
0
u/Mysterious-Effect146 3h ago
You would ever trust benchmarks by the company with a vested interest in selling you the model?
96
u/Raheeper 4h ago
from 24.7 to 52.6, what the fuck are they feeding these models
198
u/vrnvorona 4h ago
Benchmaxx fuel
•
u/Key_Reading_9664 1h ago
Terminal-Bench-Science 0.1 was released last week (as was Terminal-Bench 4.0) - https://www.tbench.ai/benchmarks
Are we saying that Fable 5.1 and Opus 5 were trained on unreleased benchmarks?
30
u/Neurogence 4h ago
How did it only improve by 2-3% in humanity's last exam?
18
u/FateOfMuffins 4h ago
Franky I'd be surprised at ever increasing scores in HLE given what we know about how many errors it likely has
0
u/kaityl3 ASI▪️2024-2027 3h ago
Oh, what do you mean?
11
u/FateOfMuffins 3h ago
Basically all benchmarks have errors in them, possibly with the exception of certain exams that thousands of people take (and even then some errors pop up every now and then, but ofc only caught because thousands of people write them).
Epoch previously estimated GPQA to have an error rate of ~7% (give or take a bit), so scores that were reaching like 94% was already a bit sus (or perhaps just signaling that effectively 100% of the benchmark was solved). They also estimated the same for Frontier Math... only for a HUGE amount of errors to be found in Frontier Math (we're talking about scores jumping from 40% to 80% level of errors)
Some people have audited some sections of HLE and claimed somewhere around (I don't remember the exact numbers) 1/3 to 1/2 of Chemistry problems had flaws and frankly idk how much that extends to the rest of the benchmark. So actual max score of HLE is unknown but I wouldn't be surprised if like 1/3 of the benchmark had errors and we're near the cap.
12
u/Long_comment_san 4h ago
That one is not open source
3
u/krzonkalla 4h ago
it is, actually, very much opensource
-1
u/Long_comment_san 3h ago
hmm, must have mistaken it for something else. Is artificial analysis intelligence closed source?
1
u/throwaway131072 2h ago
Not closed source, just supposed to have some public trials and some hidden trials that are never supposed to be leaked, and become invalid as soon as they are.
•
2
u/DelphiTsar 3h ago
Even if you are in the field, you'd still probably get most of the questions in that field wrong. It was also basically built around what AI's can't do. If LLM's were answering the question reliably they rejected the submitted question.
2
u/Saedeas 3h ago
Honestly, they're probably starting to saturate that benchmark. There are a shitton of errors in most of these top tier benchmarks, kinda by their very nature. HLE in particular is a ton of incredibly niche, expert questions, which makes it really tough to validate for correctness.
This can make improvements look weird. If for example, there's a hard cap of 70% in scoring (due to say 30% errors across the benchmark), 2% from 63-> 65 represents a 28% reduction in errors.
1
u/nothis AGI by 2030 but we'll be disappointed 2h ago edited 2h ago
They have a pretty good website, people should refer to it more often rather than treating it as some mystery: https://lastexam.ai/
I'm honestly a bit underwhelmed by the set of questions. The "hardness" seems to lie in more and more obscure domain-knowledge being necessary to answer them. Things that might not be ready available on "the internet", buried in niche publications, researcher communication that remains unpublished. So there's less of it in the training data. It's not logic-based as is ARC-AGI (which has its own problems, IMO, kinda the opposite). If they dug up the answers in obscure textbooks, great, if not, the models won't magically "reason" into existence some name or fact from a niche science.
•
u/Saedeas 1h ago
I'm familiar with the benchmark, I'm just pointing out how many errors are present in the baseline version of it.
There's a really good paper from a couple weeks ago where they characterize it: https://arxiv.org/html/2602.13964v4
Models see huge accuracy gains once you actually fix the benchmark (7-10 points absolute and 30-40 points on the questions with erroneous statements or answers).
3
1
u/FatPsychopathicWives 3h ago
That benchmark has been going on since January 2025. It's a tough one.
26
u/BarisSayit 4h ago
Major jumps in a few benchmarks are expected, one must look at the average. That's why AA score matters more.
3
7
u/voyt_eck 4h ago
Basically feeding with benchmarks and doing benchmarkmaxxing.
1
u/Healthy-Nebula-3603 3h ago
Your knowledge seems from 2025.
They just using RL now ( the model is self improving by thinking a lot during RL )
4
u/Ormusn2o 4h ago
I had same feeling using 5.6, like they put something weird in the water while training it, because coding with it has been so much substantially better than with 5.5.
Fable 5.1 being this good might indicate that at least Anthropic, figured something out on how to actually release those models, because both Mythos/Fable and gpt 5.5/5.6 started effectively around february, and I think most of the year it was both companies struggling in not releasing extremely unsafe models, so maybe this is a sign that they can finally release them safely.
2
1
u/Bitsquire 2h ago
Pretty easy to do these days tbh on benchmarks where LLMs have little targeted training data for. RLing with even just a few thousand prompts can yield massive gains.
1
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 4h ago
the models are hungry for synthetic data :3
-1
15
36
u/reefine 4h ago
Gentle reminder that Opus 5 benchmarks higher than Fable 5 in some categories and we know how that has gone.
7
u/Ok_Display_3159 4h ago
I don't use Anthropic models much, what happened with Opus 5?
7
u/Neurogence 4h ago
The main complaint is that Opus 5 is unintelligible because it's too wordy and uses overly complicated words.
One of the main highlights for Fable 5.1 is that it writes in "Plain English."
6
u/Vaughn 4h ago
Opus 5 is certainly wordy, but I wouldn't call it "unintelligible". Its writing is quite clear; it just refuses to make any assumptions about what I might know. Unless I tell it to.
•
u/LeastCounterculture 1h ago
when i use it, basic concepts start getting reduced to terms only the llm uses
and like it just keeps doing that reduction but for everything
until at the end i need to basically have it describe wtf it just said.
-1
u/Pls-No-Bully 3h ago
It’s like working with a junior dev who refuses to do any of its own research. It’s like it panics if you don’t tell it exactly where to look for everything… it’s the only recent model that has struggled with my company’s monorepo
1
•
u/Key_Reading_9664 29m ago
read: a large percentage of the folks that post on these subreddits read at an 8th grade level
11
u/Temporary-Paper5202 4h ago
Nothing, it's a good model some people are just whiny bitches who can't prompt it properly.
7
u/Foreign_Skill_6628 3h ago
There is a load-bearing assumption hidden in that comment, cross-validated by the decision record and flagged prior in the pre-planning workflow. I wanted to make you aware of this because an ignored flag is just that, a stale reminder of poor engineering discipline that should be fixed up-front, not reworked behind-the-scenes. I can correct that for you, just say go.
———
That 👆🏼is what is wrong with Opus-5. It comments, codes, and talks like an LLM mind-virus infected it and is slowly rotting its inference weights.
•
u/ultio 1h ago edited 1h ago
I had a problem with virtual machine freezes and asked Opus 5 to help me with it and it gave me an insanely convoluted response, then I asked it to shorten and simplify and it gave me this:
[...] The resize killed it because on VMSVGA a resolution change means tearing down the screen target and waiting on a fence. Miss that signal and the driver blocks forever, taking fbcon with it — the guest keeps running, you just lose the display. And it died at local-premount because Ubuntu ships KMS drivers in the initramfs for plymouth, so vmwgfx loads and wedges right there. That script wasn't the problem, it was just the last thing to reach a live framebuffer.
Now, I never gave any indication that I have intricate knowledge into Linux or acted like a smart-ass towards it, so throwing around phrases like "Ubuntu ships KMS drivers in the initramfs for plymouth" really made me laugh because that really could just be made-up techno-babble.
I asked it why it keeps talking like this and it basically told me "well I thought you're a software developer so I assume you know all of this", as if I was a freaking Linux kernel-level developer with Linus Torvalds' brain. Basically the AI version of "Oh you think you're a real gamer? Name every game!".
1
6
u/Pls-No-Bully 3h ago
Are you a professional SWE working in a massive monorepo with over a thousand other engineers?
If so, you’ll know that Opus 5 is garbage. My prompts work fine for Fable, Opus 4.6-4.8, Sol, and even Kimi when our company trialed it. Opus 5 begins to panic if you leave any ambiguity at all for a straight-forward task it should be easily capable of figuring out itself (and which all other models can easily figure out). It’s a horrible model relative to the rest
9
u/86784273 3h ago
I'm a professional SWE, work in large codebases all the time, never had issues with O5. I dont work with a thousand other engineers in the same codebase though. What counts as a massive monorepo to you? Thousands of files and millions of lines?
6
u/Temporary-Paper5202 3h ago
> Are you a professional SWE working in a massive monorepo with over a thousand other engineers?
Yes, again, skill issue.
-1
3
u/TheCraxo 4h ago
Barely used it but something similar to Sonnet 5, it gets things done after spending millions of tokens and overthinking and being corrected multiple times.
4
u/Cubewood 4h ago
A lot of the people using Opus are Vibe coders who don't actually understand what they are building. Opus 5 treats you like an equal who understands what they are working on, and this freezes many people's brains because these models have far surpassed their own capabilities at this point.
4
u/Pls-No-Bully 3h ago
Lmao this might be the worst take ever. If you leave any ambiguity for Opus 5, it begins to panic and spiral unlike any of the other models (including Opus 4.6-4.8)
I’m at a FAANG equivalent and nobody I know uses Opus 5, everyone uses Fable and/or Sol. You can’t trust it to do anything unless you hold its hand through every little step, whereas Fable and anything after Opus 4.6 use tools far more effectively. It’s like babysitting
•
u/SOCSChamp 58m ago
Not the criticism I typically see here, which I also share. The problem is that talking to it is difficult. It explains everything in roundabout ways using ridiculous analogies that have nothing to do with the context and generally reads like slop. If I'm coding, I don't want it to come back to me and say that something is load bearing, we were rolling two dice but kept one frozen, that we fired multiple shots at the right target with the wrong gun, or any of the other stupid analogies it strings together. I'm perfectly capable of understanding the subject matter but it communicates terribly. Its also noticeably worse at understanding intent or common sense than fable.
•
u/Cubewood 42m ago
Don't really have any problems with this, yes it maybe overly explains what it is doing, but I much rather have more information than too little information.
I also like that it pushes back when it believes there might be a better way of doing something, we have seen in the past what happens when these AI's just agree with everything you say, and I am not too proud to admit that an AI may know how to implement something better than me. If you still disagree with its suggestion, you can street it in a different direction and it will follow through with your suggestion.
Of course Fable is better, but Fable is also very expensive and burns through your tokens, so I only use this for very complex implementations, and primarily for security reviews, since I am on the Enterprise Plan and only get $1000,- of tokens each month.
0
14
u/reddit_guy666 4h ago edited 4h ago
What does partial computer use mean? Like human intervention needed in between?
12
u/winless 3h ago
Strict: task scoring is binary, either the model passes by satisfying every single requirement of the task or it fails.
Partial: the model can earn partial credit for correctly reaching intermediate task checkpoints (there's an average of 27.25 checkpoints per task), even if they don't ultimately meet all of the task's requirements.
-9
u/riqvip 4h ago
I think they’re just making shit up to make the model look cooler
15
u/Saedeas 4h ago
If only there were some way to look this up instead of instantly defaulting to braindead skepticism. Something like googling the name of the benchmark in question (OSWorld 2.0) and "partial scoring".
Maybe you'd find out it means something like measuring an AI agent's progress by grading smaller checkpoints along a long task instead of using a simple pass-or-fail grade.
Alas, instead we just have to be dismissive!
4
u/No_Most_5528 4h ago
Can someone explain to me how tf they jump from 20 percent to percent?
•
u/Hans-Wermhatt 1h ago
AI research is now heavily targeting science. They are shifting into scientific workflows and obviously that's why that benchmark stands out. They are a company and they identified drug discovery and science as potentially lucrative fields that are also the next verifiable domain. Expect big agentic scientific improvements in GPT as well. They are just spending a lot of time now sculpting how the model reasons in a scientific workflow.
3
u/Sunstorm84 2h ago
Benchmaxxing
•
u/Key_Reading_9664 1h ago edited 25m ago
That particular benchmark was released last week - https://www.tbench.ai/benchmarks.
Seeing people claiming that Opus 5 must be benchmaxxed because of its Terminal-Bench 4.0 score. Terminal-Bench 4.0 was also released last week
39
u/SonOfThomasWayne 4h ago
Opus 5 is hot garbage of a model and was better in all benchmarks compared to fable. I don't believe these numbers at all
8
u/kaityl3 ASI▪️2024-2027 3h ago
IDK I use Opus 5 all the time and they work great; they're just overly verbose and compulsively over-test everything. But their work is fine.
I've had the best luck with Fable as a coordinator of a swarm of Opus 5 subagents though
•
u/Minimonium 1h ago
"Great" is relative.
Even in subagent setups Opus is a bit homeless because Fable is just better at everything, just tune effort down for intermediate tasks. Second reviewer maybe, but might as well use any of the Chinese frontier models and chances are they'd spot more issues than Opus 5.
23
u/Neurogence 4h ago
Benchmarks are useless now.
The only real benchmark is employment rate and new scientific discoveries.
6
3
u/zoomoutalot 3h ago
I don't know it it was a deliberate choice but I like how you used "employment rate" and not "unemployment rate"
0
2
u/Specialist_Dark_3668 4h ago
Another reason to believe even Anthropic doesn't trust these benchmarks is that Anthropic doesn't put similar safeguards on Opus 5 even though it is benchmarked as supposedly smarter than Fable.
2
u/KaMaFour 4h ago
Aside from terminal bench science (first time i see this benchmark... ever) looks like a next step after opus 5. I'm not saying this is a bad thing though.
2
u/Work_Owl 3h ago
I feel like whenever there's a release all the benchmarks look good but it never translate into solid gains for my actual work tasks. Fable and Opus 5 are still dogshit at customer facing analysis and presenting data - yeah it can produce charts and tables from data, but the language used in reports is just weird. I also can't trust it to make logical analytical decisions like when to use averages or display all records in reports
4
u/Reddit_User_Original 3h ago
Benchmark high in science when it will refuse to talk to you about science 😂
5
u/OneConfident7361 4h ago
isn't opus 5; 2 times cheaper? considering that, trust me bro benchmark doesn't seem that impressive
2
2
u/TheSwordItself 4h ago
Mega guard rails on this thing, way worse than fable 5, can't discuss chemistry practically at all without an opus 5 swap
2
2
u/ezjakes 4h ago
Decent, but no massive jump.
0
u/sunstersun 3h ago
like they described, it's a cool improvement. nothing earth shattering like Mythos was.
I have more hopes for GPT6 and Doug. Seems like OpenAI is getting an edge due to computing available for training.
1
1
1
1
u/Error_404_403 3h ago
I am not sure how relevant those comparisons are when actual performance of same Opus model on similar task can substantially differ depending on the release date and time of day even. It looks like realistic performance is mostly driven by the compute that can be allocated to a task and not by even model name.
1
u/Happy_Guitar3521 3h ago
Those benchmarks just saved me another ~$50k in hardware. Experimenting hard.
1
u/Chesstiger2612 2h ago
What do you mean by that? You can do the same thing you would otherwise need to spend 50k on?
1
1
u/Mysterious-Effect146 3h ago
Every benchmark by the very company announcing it is an ad and should be taken with a mountain of salt. It is an ad.
1
u/burritos4jesus 2h ago edited 2h ago
Discussion around AI benchmarks is ALWAYS dominated by agentic coding and software engineering. Which makes sense if you're a developer, but I'm not technical. I'm not a coder. I work in sales, in an office, and I care way more about AutomationBench than I do about whether the newest model got another five points better at writing code.
My workday is spread across a bunch of biz apps. CRM, email, calendar, Teams/Zoom, shared files, internal knowledge, finance systems, sales enablement tools, etc. What I want to know is whether I can put an AI in the middle of that stack and say: here's the customer meeting, figure out what happened, pull whatever additional context you need, update the right account and opportunity, create the proper next steps, schedule the follow-up with the right people, send the appropriate templated email, and leave all of the underlying systems correct when you're finished.
That's basically what AutomationBench is trying to measure.
It only came out in April, and when Zapier launched it, the best frontier models were below 10% success. We're already around 30% a few months later depending on the model/evaluation, which is an enormous improvement, but 30% is still obviously nowhere close to "give the AI access to Salesforce and let it run unsupervised."
The scoring is also brutal but in a useful way. It isn't asking whether the model mostly understood the assignment. It checks the final state of the business systems. If the AI correctly updates four things but misses the fifth, contacts the wrong person, creates a duplicate record or says it finished when it didn't, the workflow fails. That's exactly how I WANT a benchmark graded.
There's another caveat too: AutomationBench isn't giving these models an optimized company-specific harness setup. In its evaluation, the model has to search through hundreds of possible API endpoints, figure out which tools and data it needs, execute the calls and verify the outcome by itself. A real deployment could have a much better harness around the model: pre-mapped CRM actions, company-specific rules, known account IDs, structured transcript extraction, retries, write verification, confidence thresholds and human approval for ambiguous actions.
So I don't think 30% means AI can only do 30% of my job, rather I think it means we're still pretty bad at handing a general-purpose model a messy business environment and saying "handle this entire workflow perfectly with no supervision."
But THAT is the benchmark progression I care about.
If AutomationBench goes from <10% in April 2026 to ~30% now, I want to see where it is in six months, a year, two years. Because when these models start hitting 70% on strict end-to-end business workflows, especially once paired with good harnesses, that has a much more direct impact on my working life than another record on a coding benchmark.
1
•
•
u/Fragrant-Job-3200 AGI 2026 ASI 2028 1h ago
I wonder how good would Fable 5.1 and Astra perform on ARC-AGI 3.
•
u/Present-Motor-173 40m ago
Can you use it for biology yet? Sol is far better than any Opus model for molecular biology right now. But I don't want to resubscribe unless I can use it.
1
0
u/firaristt 3h ago
This has little to no value to me. Because it's too expensive for what it can do. For really hard tasks, I have to take the wheel, for others, smaller, cheaper models do %90 of the job for a fraction of the cost. Like, literally, instead of 10-30$, they cost just a few, 2-5$.
For me unless it's emergency, this doesn't make sense. I can use same amount of tokens $$$ for 1 task with these huge and expensive models or I can work for a week on multiple tasks for the same tokens $$$.
Oh, also consider when these models can't one-shot, all the tokens go straight to the garbage bin. Whereas cheap smaller models, they won't hurt, just ask a retry.
-1
-1
u/bladerskb 4h ago
where are the rest of the benchmarks? they think we're stupid
1
u/Healthy-Nebula-3603 3h ago
Older tests are just redundant nowadays mostly. Over 90% on them has 0 sense showing it
1
u/bladerskb 3h ago
deepswe 1.1 is old and redundant?
1
u/Healthy-Nebula-3603 3h ago
I said mostly not all .
At least you have a terminal bench 4
DeepSWE 1.1 opus has 74% so this one has probably over 80%... seems almost saturated.
-1
u/Efficient-Cat-1591 3h ago
If the benchmarks are true nothing in the market beats Fable 5.1 for performance vs quality. GPT is lagging behind now, only benefit is the generous usage limits.
2
u/Supermax64 2h ago
They're days away from releasing a new model. Lagging behind is a weird take.
0
u/Efficient-Cat-1591 2h ago
days away? Evidence? Benchmarks?
•
u/steny007 1h ago
Release days away imply benchmarks days away too. That's a common knowledge even among the less gifted here.
•
290
u/LiquidNeat 4h ago