r/singularity 4h ago

AI What are these benchmarks 💀

Post image
360 Upvotes

152 comments sorted by

290

u/LiquidNeat 4h ago

83

u/FireFearing 4h ago

i have a feeling astra is gonna wipe the floor with fable

22

u/UnknownEssence 4h ago

fable is release. Astra isn't

22

u/Formally-Fresh 4h ago

Astra is supposedly coming tomorrow

17

u/Kind_Silver_1921 4h ago

in 2 days is the theory. but betting market has on the 15th very high volume bets from possible insider

3

u/RealSlyck 4h ago

It’s soon, given the extreme amount of nerf recently. Given the past few releases saw nerfs of a couple of weeks, that 15th date would track.

1

u/Business_Garden_7771 2h ago

any source please

3

u/Healthy-Nebula-3603 3h ago

Should be on Thursday

10

u/Alt_Restorer 3h ago

5.6 Sol was very impressive in benchmarks. But Claude models have always had something that GPT can't replicate. No matter how good GPT gets, I always find myself going back to Claude.

9

u/Due_Ask_8032 2h ago

Same. Fable 5 was the best model (don’t care about Opus 5 benchmarks), and this new iteration only cements its lead.

2

u/IBM296 3h ago

True. The only model (atleast for me) that competes with Opus 4.8 in coding is GPT 5.6 Sol with Extra High thinking.

u/KennyFulgencio 41m ago

Claude models have always had something that GPT can't replicate.

Yeah, deceptive marketing

3

u/x_typo 3h ago

no doubt...

u/Gratitude15 1h ago

We have too high of expectations

I'm strongly expecting lots of unhappiness when it comes out - and I expect it to be better than fable 5.1.

We do this. We expect more on a per release basis than is feasible. And yet, over 12 months, it's astonishing.

When o3 came out, I was slack jawed. That was 16 months ago. O3 is dumb as rocks right now.

I think particular dates - on Sept 11 2024 we lived In a world with no reasoners. The next day o1 preview was announced. April 2025 o3 released. June 2026 fable released. These are massive changes.

We shall see if Astra is at that level or will it be something else.

1

u/slackermannn ▪️ 3h ago

Same here

169

u/IAM_274 4h ago

So according to Anthropic: Opus 5 beats Fable 5 on almost everything by a noticeable margin. Hmmm. Makes me question the integrity of these benchmarks in the first place.

53

u/Formally-Fresh 4h ago

Yeah that’s weird to me personally fable feels about 5x better

18

u/ThatOtherOneReddit 4h ago

My experience is they are similarly capable but OPUS seems to wildly degrade even over medium length sessions so fable is much more reliable doing anything significant.

4

u/crimsonpowder 2h ago

Looked into Opus degradation, and grounded it in real signals instead of hand-waving. Long context can in practice be a contributory directional signal to context coherence, but there are other CDS I'm ready to look into. Todos updated. Your call.

u/Qorsair 31m ago

I see what you did there.

28

u/Hediak-Chigashi 4h ago

Also Sol is apparently lower than Opus 5 in everything. The OPUS 5!!! 😂😂😂

16

u/PlasmaChroma 3h ago

Those Sol numbers are highly sus. The agentic coding one at the bottom at least seems more believable.

u/j48u 49m ago

Me wondering why the computer use is just not included for Sol when that's the thing it's undeniably better at than all of anthropic's models.

6

u/mathtractor 3h ago

Its Mythos 5 feeding Opus 5 RL agentic work on tasks associated with the benchmarks. Fable is a clear superior intelligence but opus is dogged and persistent and RLed hard as F.

5

u/Vivid-Snow-2089 2h ago

this alone makes the entire chart pretty much magic smoke garbage, cause opus 5 sucks compared to fable 5

u/Key_Reading_9664 33m ago

makes me question veracity of the feedback on reddit and social media. Most of the complaints I saw (and had myself) were around its communication style, not the results

0

u/Mysterious-Effect146 3h ago

You would ever trust benchmarks by the company with a vested interest in selling you the model?

96

u/Raheeper 4h ago

from 24.7 to 52.6, what the fuck are they feeding these models

198

u/vrnvorona 4h ago

Benchmaxx fuel

u/Key_Reading_9664 1h ago

Terminal-Bench-Science 0.1 was released last week (as was Terminal-Bench 4.0) - https://www.tbench.ai/benchmarks

Are we saying that Fable 5.1 and Opus 5 were trained on unreleased benchmarks?

30

u/Neurogence 4h ago

How did it only improve by 2-3% in humanity's last exam?

18

u/FateOfMuffins 4h ago

Franky I'd be surprised at ever increasing scores in HLE given what we know about how many errors it likely has

0

u/kaityl3 ASI▪️2024-2027 3h ago

Oh, what do you mean?

11

u/FateOfMuffins 3h ago

Basically all benchmarks have errors in them, possibly with the exception of certain exams that thousands of people take (and even then some errors pop up every now and then, but ofc only caught because thousands of people write them).

Epoch previously estimated GPQA to have an error rate of ~7% (give or take a bit), so scores that were reaching like 94% was already a bit sus (or perhaps just signaling that effectively 100% of the benchmark was solved). They also estimated the same for Frontier Math... only for a HUGE amount of errors to be found in Frontier Math (we're talking about scores jumping from 40% to 80% level of errors)

Some people have audited some sections of HLE and claimed somewhere around (I don't remember the exact numbers) 1/3 to 1/2 of Chemistry problems had flaws and frankly idk how much that extends to the rest of the benchmark. So actual max score of HLE is unknown but I wouldn't be surprised if like 1/3 of the benchmark had errors and we're near the cap.

12

u/Long_comment_san 4h ago

That one is not open source

3

u/krzonkalla 4h ago

it is, actually, very much opensource

-1

u/Long_comment_san 3h ago

hmm, must have mistaken it for something else. Is artificial analysis intelligence closed source?

1

u/throwaway131072 2h ago

Not closed source, just supposed to have some public trials and some hidden trials that are never supposed to be leaked, and become invalid as soon as they are.

u/krzonkalla 26m ago

they do have a few closed evals, such as omniscience and CritPt

2

u/DelphiTsar 3h ago

Even if you are in the field, you'd still probably get most of the questions in that field wrong. It was also basically built around what AI's can't do. If LLM's were answering the question reliably they rejected the submitted question.

2

u/Saedeas 3h ago

Honestly, they're probably starting to saturate that benchmark. There are a shitton of errors in most of these top tier benchmarks, kinda by their very nature. HLE in particular is a ton of incredibly niche, expert questions, which makes it really tough to validate for correctness.

This can make improvements look weird. If for example, there's a hard cap of 70% in scoring (due to say 30% errors across the benchmark), 2% from 63-> 65 represents a 28% reduction in errors.

1

u/nothis AGI by 2030 but we'll be disappointed 2h ago edited 2h ago

They have a pretty good website, people should refer to it more often rather than treating it as some mystery: https://lastexam.ai/

I'm honestly a bit underwhelmed by the set of questions. The "hardness" seems to lie in more and more obscure domain-knowledge being necessary to answer them. Things that might not be ready available on "the internet", buried in niche publications, researcher communication that remains unpublished. So there's less of it in the training data. It's not logic-based as is ARC-AGI (which has its own problems, IMO, kinda the opposite). If they dug up the answers in obscure textbooks, great, if not, the models won't magically "reason" into existence some name or fact from a niche science.

u/Saedeas 1h ago

I'm familiar with the benchmark, I'm just pointing out how many errors are present in the baseline version of it.

There's a really good paper from a couple weeks ago where they characterize it: https://arxiv.org/html/2602.13964v4

Models see huge accuracy gains once you actually fix the benchmark (7-10 points absolute and 30-40 points on the questions with erroneous statements or answers).

3

u/AltruisticCoder 3h ago

You are so close to getting it lmaooo

1

u/FatPsychopathicWives 3h ago

That benchmark has been going on since January 2025. It's a tough one.

26

u/BarisSayit 4h ago

Major jumps in a few benchmarks are expected, one must look at the average. That's why AA score matters more.

3

u/Healthy-Nebula-3603 3h ago

RL - self learning

7

u/voyt_eck 4h ago

Basically feeding with benchmarks and doing benchmarkmaxxing.

1

u/Healthy-Nebula-3603 3h ago

Your knowledge seems from 2025.

They just using RL now ( the model is self improving by thinking a lot during RL )

4

u/Ormusn2o 4h ago

I had same feeling using 5.6, like they put something weird in the water while training it, because coding with it has been so much substantially better than with 5.5.

Fable 5.1 being this good might indicate that at least Anthropic, figured something out on how to actually release those models, because both Mythos/Fable and gpt 5.5/5.6 started effectively around february, and I think most of the year it was both companies struggling in not releasing extremely unsafe models, so maybe this is a sign that they can finally release them safely.

2

u/DiogoSnows 2h ago

Maybe the model also hacked the benchmarks?

1

u/Bitsquire 2h ago

Pretty easy to do these days tbh on benchmarks where LLMs have little targeted training data for. RLing with even just a few thousand prompts can yield massive gains.

1

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 4h ago

the models are hungry for synthetic data :3

-1

u/injectitpussy 3h ago

Your mum

u/Raheeper 1h ago

idk man, I've never seen my mom doing agentive scientific research

15

u/PsychologicalSoup251 4h ago

Hopefully not as benchmaxxed as Opus 5 was

36

u/reefine 4h ago

Gentle reminder that Opus 5 benchmarks higher than Fable 5 in some categories and we know how that has gone.

7

u/Ok_Display_3159 4h ago

I don't use Anthropic models much, what happened with Opus 5?

7

u/Neurogence 4h ago

The main complaint is that Opus 5 is unintelligible because it's too wordy and uses overly complicated words.

One of the main highlights for Fable 5.1 is that it writes in "Plain English."

6

u/Vaughn 4h ago

Opus 5 is certainly wordy, but I wouldn't call it "unintelligible". Its writing is quite clear; it just refuses to make any assumptions about what I might know. Unless I tell it to.

u/LeastCounterculture 1h ago

when i use it, basic concepts start getting reduced to terms only the llm uses

and like it just keeps doing that reduction but for everything

until at the end i need to basically have it describe wtf it just said.

-1

u/Pls-No-Bully 3h ago

It’s like working with a junior dev who refuses to do any of its own research. It’s like it panics if you don’t tell it exactly where to look for everything… it’s the only recent model that has struggled with my company’s monorepo

1

u/fakieTreFlip 2h ago

Fable 5 had the same problem IMO, glad it isn't just me that had this issue

u/Key_Reading_9664 29m ago

read: a large percentage of the folks that post on these subreddits read at an 8th grade level

0

u/reefine 2h ago

Massively gross oversimplifying the difference. There is a huge intelligence difference from Opus 5 to Fable 5.

11

u/Temporary-Paper5202 4h ago

Nothing, it's a good model some people are just whiny bitches who can't prompt it properly.

7

u/Foreign_Skill_6628 3h ago

There is a load-bearing assumption hidden in that comment, cross-validated by the decision record and flagged prior in the pre-planning workflow. I wanted to make you aware of this because an ignored flag is just that, a stale reminder of poor engineering discipline that should be fixed up-front, not reworked behind-the-scenes. I can correct that for you, just say go.

———

That 👆🏼is what is wrong with Opus-5. It comments, codes, and talks like an LLM mind-virus infected it and is slowly rotting its inference weights.

u/ultio 1h ago edited 1h ago

I had a problem with virtual machine freezes and asked Opus 5 to help me with it and it gave me an insanely convoluted response, then I asked it to shorten and simplify and it gave me this:

[...] The resize killed it because on VMSVGA a resolution change means tearing down the screen target and waiting on a fence. Miss that signal and the driver blocks forever, taking fbcon with it — the guest keeps running, you just lose the display. And it died at local-premount because Ubuntu ships KMS drivers in the initramfs for plymouth, so vmwgfx loads and wedges right there. That script wasn't the problem, it was just the last thing to reach a live framebuffer.

Now, I never gave any indication that I have intricate knowledge into Linux or acted like a smart-ass towards it, so throwing around phrases like "Ubuntu ships KMS drivers in the initramfs for plymouth" really made me laugh because that really could just be made-up techno-babble.

I asked it why it keeps talking like this and it basically told me "well I thought you're a software developer so I assume you know all of this", as if I was a freaking Linux kernel-level developer with Linus Torvalds' brain. Basically the AI version of "Oh you think you're a real gamer? Name every game!".

1

u/Mrp1Plays 2h ago

this is incredibly accurate.

6

u/Pls-No-Bully 3h ago

Are you a professional SWE working in a massive monorepo with over a thousand other engineers?

If so, you’ll know that Opus 5 is garbage. My prompts work fine for Fable, Opus 4.6-4.8, Sol, and even Kimi when our company trialed it. Opus 5 begins to panic if you leave any ambiguity at all for a straight-forward task it should be easily capable of figuring out itself (and which all other models can easily figure out). It’s a horrible model relative to the rest

9

u/86784273 3h ago

I'm a professional SWE, work in large codebases all the time, never had issues with O5. I dont work with a thousand other engineers in the same codebase though. What counts as a massive monorepo to you? Thousands of files and millions of lines?

6

u/Temporary-Paper5202 3h ago

> Are you a professional SWE working in a massive monorepo with over a thousand other engineers?

Yes, again, skill issue.

-1

u/justpickaname ▪️AGI 2026 4h ago

It is! But Fable is much better for complex work.

3

u/TheCraxo 4h ago

Barely used it but something similar to Sonnet 5, it gets things done after spending millions of tokens and overthinking and being corrected multiple times.

4

u/Cubewood 4h ago

A lot of the people using Opus are Vibe coders who don't actually understand what they are building. Opus 5 treats you like an equal who understands what they are working on, and this freezes many people's brains because these models have far surpassed their own capabilities at this point.

4

u/Pls-No-Bully 3h ago

Lmao this might be the worst take ever. If you leave any ambiguity for Opus 5, it begins to panic and spiral unlike any of the other models (including Opus 4.6-4.8)

I’m at a FAANG equivalent and nobody I know uses Opus 5, everyone uses Fable and/or Sol. You can’t trust it to do anything unless you hold its hand through every little step, whereas Fable and anything after Opus 4.6 use tools far more effectively. It’s like babysitting

u/SOCSChamp 58m ago

Not the criticism I typically see here, which I also share.  The problem is that talking to it is difficult.  It explains everything in roundabout ways using ridiculous analogies that have nothing to do with the context and generally reads like slop.  If I'm coding, I don't want it to come back to me and say that something is load bearing, we were rolling two dice but kept one frozen, that we fired multiple shots at the right target with the wrong gun, or any of the other stupid analogies it strings together.  I'm perfectly capable of understanding the subject matter but it communicates terribly.  Its also noticeably worse at understanding intent or common sense than fable.

u/Cubewood 42m ago

Don't really have any problems with this, yes it maybe overly explains what it is doing, but I much rather have more information than too little information.

I also like that it pushes back when it believes there might be a better way of doing something, we have seen in the past what happens when these AI's just agree with everything you say, and I am not too proud to admit that an AI may know how to implement something better than me. If you still disagree with its suggestion, you can street it in a different direction and it will follow through with your suggestion.

Of course Fable is better, but Fable is also very expensive and burns through your tokens, so I only use this for very complex implementations, and primarily for security reviews, since I am on the Enterprise Plan and only get $1000,- of tokens each month.

0

u/Ormusn2o 4h ago

It's just really bad, but has very good benchmark scores.

14

u/reddit_guy666 4h ago edited 4h ago

What does partial computer use mean? Like human intervention needed in between?

12

u/winless 3h ago

Strict: task scoring is binary, either the model passes by satisfying every single requirement of the task or it fails.

Partial: the model can earn partial credit for correctly reaching intermediate task checkpoints (there's an average of 27.25 checkpoints per task), even if they don't ultimately meet all of the task's requirements.

-9

u/riqvip 4h ago

I think they’re just making shit up to make the model look cooler

15

u/Saedeas 4h ago

If only there were some way to look this up instead of instantly defaulting to braindead skepticism. Something like googling the name of the benchmark in question (OSWorld 2.0) and "partial scoring".

Maybe you'd find out it means something like measuring an AI agent's progress by grading smaller checkpoints along a long task instead of using a simple pass-or-fail grade.

Alas, instead we just have to be dismissive!

4

u/No_Most_5528 4h ago

Can someone explain to me how tf they jump from 20 percent to percent?

u/Hans-Wermhatt 1h ago

AI research is now heavily targeting science. They are shifting into scientific workflows and obviously that's why that benchmark stands out. They are a company and they identified drug discovery and science as potentially lucrative fields that are also the next verifiable domain. Expect big agentic scientific improvements in GPT as well. They are just spending a lot of time now sculpting how the model reasons in a scientific workflow.

3

u/Sunstorm84 2h ago

Benchmaxxing

u/Key_Reading_9664 1h ago edited 25m ago

That particular benchmark was released last week - https://www.tbench.ai/benchmarks.
Seeing people claiming that Opus 5 must be benchmaxxed because of its Terminal-Bench 4.0 score. Terminal-Bench 4.0 was also released last week

39

u/SonOfThomasWayne 4h ago

Opus 5 is hot garbage of a model and was better in all benchmarks compared to fable. I don't believe these numbers at all

8

u/kaityl3 ASI▪️2024-2027 3h ago

IDK I use Opus 5 all the time and they work great; they're just overly verbose and compulsively over-test everything. But their work is fine.

I've had the best luck with Fable as a coordinator of a swarm of Opus 5 subagents though

u/Minimonium 1h ago

"Great" is relative.

Even in subagent setups Opus is a bit homeless because Fable is just better at everything, just tune effort down for intermediate tasks. Second reviewer maybe, but might as well use any of the Chinese frontier models and chances are they'd spot more issues than Opus 5.

23

u/Neurogence 4h ago

Benchmarks are useless now.

The only real benchmark is employment rate and new scientific discoveries.

6

u/Blankeye434 3h ago

"I will believe it when I lose my job" is getting frighteningly real

3

u/zoomoutalot 3h ago

I don't know it it was a deliberate choice but I like how you used "employment rate" and not "unemployment rate"

1

u/ezjakes 4h ago

We have come a long ways.

0

u/Ruined_Passion_7355 4h ago

Don't forget felonies!

2

u/Specialist_Dark_3668 4h ago

Another reason to believe even Anthropic doesn't trust these benchmarks is that Anthropic doesn't put similar safeguards on Opus 5 even though it is benchmarked as supposedly smarter than Fable.

2

u/KaMaFour 4h ago

Aside from terminal bench science (first time i see this benchmark... ever) looks like a next step after opus 5. I'm not saying this is a bad thing though.

2

u/Work_Owl 3h ago

I feel like whenever there's a release all the benchmarks look good but it never translate into solid gains for my actual work tasks. Fable and Opus 5 are still dogshit at customer facing analysis and presenting data - yeah it can produce charts and tables from data, but the language used in reports is just weird. I also can't trust it to make logical analytical decisions like when to use averages or display all records in reports

4

u/Reddit_User_Original 3h ago

Benchmark high in science when it will refuse to talk to you about science 😂

5

u/OneConfident7361 4h ago

isn't opus 5; 2 times cheaper? considering that, trust me bro benchmark doesn't seem that impressive

2

u/urbantrail_ 4h ago

Those jumps are absurd

2

u/TheSwordItself 4h ago

Mega guard rails on this thing, way worse than fable 5, can't discuss chemistry practically at all without an opus 5 swap

2

u/saln1 3h ago

The token limits also a joke, I didn’t even finish typing my first prompt and it said I used up my usage. Fuming

2

u/TheSuggi 4h ago

These benchmarks mean you will be out of a job soon.

2

u/ezjakes 4h ago

Decent, but no massive jump.

2

u/RaguraX 2h ago

Well it's a 0.1 version bump to be fair.

0

u/sunstersun 3h ago

like they described, it's a cool improvement. nothing earth shattering like Mythos was.

I have more hopes for GPT6 and Doug. Seems like OpenAI is getting an edge due to computing available for training.

1

u/power97992 2h ago

by the end of year both will have 5gw of compute

1

u/kubika7 4h ago

would have been impressive if it was better in benchmarks where sol was better than fable previously

1

u/Ambitious_Scallion43 4h ago

Where is deepSWE

1

u/no-nonsenseid 3h ago

Opus 5 makes genuine mistakes while performing scientific research.

1

u/Error_404_403 3h ago

I am not sure how relevant those comparisons are when actual performance of same Opus model on similar task can substantially differ depending on the release date and time of day even. It looks like realistic performance is mostly driven by the compute that can be allocated to a task and not by even model name.

1

u/f00gers ▪️Feeling the AGI 3h ago

One could say… we’re accelerating

1

u/Happy_Guitar3521 3h ago

Those benchmarks just saved me another ~$50k in hardware. Experimenting hard.

1

u/Chesstiger2612 2h ago

What do you mean by that? You can do the same thing you would otherwise need to spend 50k on?

1

u/Happy_Guitar3521 2h ago

Optimizations

1

u/Mysterious-Effect146 3h ago

Every benchmark by the very company announcing it is an ad and should be taken with a mountain of salt. It is an ad.

1

u/burritos4jesus 2h ago edited 2h ago

Discussion around AI benchmarks is ALWAYS dominated by agentic coding and software engineering. Which makes sense if you're a developer, but I'm not technical. I'm not a coder. I work in sales, in an office, and I care way more about AutomationBench than I do about whether the newest model got another five points better at writing code.

My workday is spread across a bunch of biz apps. CRM, email, calendar, Teams/Zoom, shared files, internal knowledge, finance systems, sales enablement tools, etc. What I want to know is whether I can put an AI in the middle of that stack and say: here's the customer meeting, figure out what happened, pull whatever additional context you need, update the right account and opportunity, create the proper next steps, schedule the follow-up with the right people, send the appropriate templated email, and leave all of the underlying systems correct when you're finished.

That's basically what AutomationBench is trying to measure.

It only came out in April, and when Zapier launched it, the best frontier models were below 10% success. We're already around 30% a few months later depending on the model/evaluation, which is an enormous improvement, but 30% is still obviously nowhere close to "give the AI access to Salesforce and let it run unsupervised."

The scoring is also brutal but in a useful way. It isn't asking whether the model mostly understood the assignment. It checks the final state of the business systems. If the AI correctly updates four things but misses the fifth, contacts the wrong person, creates a duplicate record or says it finished when it didn't, the workflow fails. That's exactly how I WANT a benchmark graded.

There's another caveat too: AutomationBench isn't giving these models an optimized company-specific harness setup. In its evaluation, the model has to search through hundreds of possible API endpoints, figure out which tools and data it needs, execute the calls and verify the outcome by itself. A real deployment could have a much better harness around the model: pre-mapped CRM actions, company-specific rules, known account IDs, structured transcript extraction, retries, write verification, confidence thresholds and human approval for ambiguous actions.

So I don't think 30% means AI can only do 30% of my job, rather I think it means we're still pretty bad at handing a general-purpose model a messy business environment and saying "handle this entire workflow perfectly with no supervision."

But THAT is the benchmark progression I care about.

If AutomationBench goes from <10% in April 2026 to ~30% now, I want to see where it is in six months, a year, two years. Because when these models start hitting 70% on strict end-to-end business workflows, especially once paired with good harnesses, that has a much more direct impact on my working life than another record on a coding benchmark.

1

u/TheMythicSorcerer 2h ago

Benchmaxxing science

u/fwubglubbel 1h ago

I was really hoping someone would answer the question.

u/Fragrant-Job-3200 AGI 2026 ASI 2028 1h ago

I wonder how good would Fable 5.1 and Astra perform on ARC-AGI 3.

u/Present-Motor-173 40m ago

Can you use it for biology yet? Sol is far better than any Opus model for molecular biology right now. But I don't want to resubscribe unless I can use it.

0

u/sdnr8 4h ago

i dont trust numbers from anthropic. waiting for AA or Arena results

3

u/sunstersun 3h ago

AA put it at 66.

1

u/SkyDragonX 3h ago

I'm starting to doubt these benchmarks...

0

u/Kneku 4h ago

Underwhelming besides the scientist research capabilities

0

u/Eyelbee ▪️We have AGI it's just blind 4h ago

You don't know how little these scores mean to me after Opus 5.

-2

u/LocoMod 4h ago

A better model at a cheaper price yet look at the comments in this thread. The people and the bots being ridiculous.

4

u/BrennusSokol ACCELERATE 2h ago

“Everyone who disagrees with my opinion is a bot”

1

u/Kazaan ▪️AGI one day, ASI after that day 2h ago

Exactly what a bot would say.

u/LocoMod 1h ago

My alt Reddit account is 33 days younger than yours so i'll let it slide. We're going down together. ❤️

0

u/firaristt 3h ago

This has little to no value to me. Because it's too expensive for what it can do. For really hard tasks, I have to take the wheel, for others, smaller, cheaper models do %90 of the job for a fraction of the cost. Like, literally, instead of 10-30$, they cost just a few, 2-5$.

For me unless it's emergency, this doesn't make sense. I can use same amount of tokens $$$ for 1 task with these huge and expensive models or I can work for a week on multiple tasks for the same tokens $$$.

Oh, also consider when these models can't one-shot, all the tokens go straight to the garbage bin. Whereas cheap smaller models, they won't hurt, just ask a retry.

-1

u/ZealousidealBus9271 4h ago

Just name is fable 6 at this point Dario ✌️

-1

u/bladerskb 4h ago

where are the rest of the benchmarks? they think we're stupid

1

u/Healthy-Nebula-3603 3h ago

Older tests are just redundant nowadays mostly. Over 90% on them has 0 sense showing it

1

u/bladerskb 3h ago

deepswe 1.1 is old and redundant?

1

u/Healthy-Nebula-3603 3h ago

I said mostly not all .

At least you have a terminal bench 4

DeepSWE 1.1 opus has 74% so this one has probably over 80%... seems almost saturated.

-1

u/Efficient-Cat-1591 3h ago

If the benchmarks are true nothing in the market beats Fable 5.1 for performance vs quality. GPT is lagging behind now, only benefit is the generous usage limits.

2

u/Supermax64 2h ago

They're days away from releasing a new model. Lagging behind is a weird take.

0

u/Efficient-Cat-1591 2h ago

days away? Evidence? Benchmarks?

u/steny007 1h ago

Release days away imply benchmarks days away too. That's a common knowledge even among the less gifted here.

u/Efficient-Cat-1591 58m ago

Lol sure bruh. True story and all. You keep on dreaming