r/OpenAI 10h ago

News Fable 5.1 released. Significant benchmark improvements, what do you think?

Same input/output pricing and 75% price reduction for cache.

211 Upvotes

89 comments sorted by

176

u/ben_bliksem 10h ago

I think we'll see countless reviews by the Reddit crowd over the next couple days, it'll be the most amazing thing and then in two weeks' time the real sentiment will surface.

Just like every other time.

32

u/liright 9h ago

Best use the crap out of the model the next few days before they nerf it to conserve compute.

2

u/Maleficent_Disk9583 1h ago

It's funny how predictable the AI companies have become. Like clockwork, hahaha

22

u/Fresh_Sock8660 9h ago

"is it me or did fable get dumber today"

An ageless classic. 

10

u/Treadmillrunner 8h ago

To be fair Anthropic does make a habit of making the models extra powerful for a week or two after release before nerfing them a bit.

2

u/interwebzdotnet 1h ago

I'm using it for a high level review on my anti Flock) ALPR project right now and it's doing a pretty good strategy, efficiency, progress, capabilities, and accuracy analysis right now.

Short run so far, but I'm happy with what I'm seeing.

2

u/ProtoplanetaryNebula 7h ago

There is no doubt the trend is upwards though.

1

u/drspock99 5h ago

That's because they nurf it bro.

1

u/Longjumping_Stop6269 9h ago

Reel em in, get em hooked, then have them beg

57

u/Bloated_Plaid 10h ago

Fucking crazy for 0.1 but the important part is lower cost on cache reads which was killing me. Cannot wait for Astra, it’s gonna be insane.

13

u/CartographerAble9446 9h ago

the cache cost doesnt matter much for average subscription, it's only for the API, and most of us are peasants and dont pay API price

3

u/Bloated_Plaid 9h ago

where did you see that the cost savings dont transfer to subs? They havent explicitly said anything like that.

3

u/hydralisk_hydrawife 7h ago

he's wrong, cache savings translates into lower costs for the company which allows more usage for the users

3

u/RealSuperdau 7h ago

Unless companies actually start using Fable/Mythos with the on-premise retention policies, then subscriptions are fucked

1

u/Seerix 6h ago

"As mentioned above, we have reduced the price of Fable 5.1’s cache reads (where the model reuses context it has already processed) wherever usage is billed by token, such as on our API"

2

u/Bolaumius 9h ago

It looks like the lowest cost on cache reads applies only for API.

1

u/ThunderStorm420 9h ago

Why do you think that Astra will be insane? I see all hype no facts, unless I don't know something

2

u/EbbExternal3544 9h ago

Because it will have to defeat Fable 5.1

1

u/Healthy-Nebula-3603 9h ago

Apparently a lot.

Look on the leaks from X

1

u/Bloated_Plaid 9h ago

Lot of outputs on Twitter.

3

u/ThunderStorm420 8h ago

I don't doubt that it will be good. But X "leaks" are of questionable authority.

2

u/Bloated_Plaid 8h ago

There is no other good source for AI discussion and discourse right now unfortunately.

0

u/coloradical5280 8h ago

Not really though, if you know who to trust they’re basically always consistent and correct.

And this always on agent that just is just, always running, forever, like a real assistant, sounds like it’s for sure being announced in dev day on the 29th

1

u/JohnToFire 7h ago

Ok who do we trust ?

1

u/HarvestingMomentum 2h ago

"leakers" with anime pfps

10

u/EbbExternal3544 9h ago

So does that mean Astra should release this week? 

17

u/ThunderStorm420 8h ago edited 8h ago

I'm sure they're hurrying now. Only thing left to wait after Astra is a 50% usage increase from Anthropic.

1

u/Your_mortal_enemy 8h ago

Everyone keeps saying that but they're releasing something massive at openai dev day at the end of the month and it's this... Surely? I hope Im wrong, almost four weeks seems like an eternity

9

u/Supermax64 8h ago

AA benchmark lists it as nearly 4x more expensive than Sol (max) to complete a task. Am I missing something?

6

u/RealSuperdau 7h ago

Damn, it's 17% more expensive on AA than Fable 5 max, despite cutting cached input rates by 4x.

Fable 5.1 max just seems really token inefficient.

0

u/reefine 5h ago

Tunnel vision. Intelligence is a factor.

4

u/Keep-Darwin-Going 7h ago

One major upgrade is it stop talking like some obnoxious tech bro that try to show off all the time. The progress and decision making feels more like normal human being

2

u/f3xjc 7h ago

Now I want an opus 5.1

11

u/Euphoric_North_745 9h ago

don't care, codex is currently working, writing code, no need to mess up my work, it is already in progress

2

u/djLiTh 7h ago

Seeing opus5 have any score higher than 5.6 makes me think this is all made up. I’m mostly doing scientific research and opus is unusable, fable5 is usable but makes mistakes but sometimes comes up with things 5.6 doesn’t, but 5.6 is currently in the lead. I’ll be trying out fable 5.1, as my workflow is to leverage two modals against each other as I find that fingers the best result. So gpt 5.6 to Fable 5.1 and back and forth.

2

u/Flaxseed4138 9h ago

How is it at coding though? And is it still 5x more token-heavy than Opus?

5

u/Key-Injury-1875 9h ago

Complete shit. Ran out of usage already on max 20x

12

u/ThunderStorm420 8h ago

What the hell are you running, genuinely?

28

u/Ryan526 8h ago

He made another to-do kanban board

2

u/Key-Injury-1875 8h ago

For Sol, it reasons much better faster and using way less usage

1

u/Key-Injury-1875 8h ago

Autonomous drone racing stack it didnt even start coding it literally just reasoned on what I should do

5

u/unfathomably_big 7h ago

To use your entire 20x with a single prompt has got to be the worst prompt engineering I’ve ever seen. This sub needs a Wall Street bets style loss porn tag for you “I only asked it to digest encyclopaedia Brittanica and now I’m out of dollars” guys

2

u/BellacosePlayer 4h ago

"pls make gta7, make no mistakes"

1

u/Nuzina 4h ago

I don’t even know how it’s possible

1

u/Practical-Positive34 9h ago

For coding I think a 3-4% improvement at the cost of nearly 2x as much is simply not worth it. I would only use this after Sonnet and Opus have coded things to go find or solve things that are broken that Opus or Sonnet couldn't fix properly. My 2 cents.

3

u/Healthy-Nebula-3603 9h ago

Did you even read ? Is cheaper a lot comparing to fable 5.

1

u/Practical-Positive34 4h ago

Making something absurdly expensive slightly cheaper does not make that absurdly thing cheaper than what it was already absurdly more expensive in comparison to. Its shocking you don't get this very basic concept.

0

u/Healthy-Nebula-3603 4h ago

I just answered to your X2 more expensive claim.

Your current answer had 0 logic and sense to the topic you mentioned.

1

u/Healthy-Nebula-3603 9h ago

We will be thinking on Thursday:)

1

u/Individual-Hunt9547 9h ago

Not sustainable at current pricing

1

u/creamyshart 8h ago

Artificial Analysis shows it costing significantly more than 5 to run its tests.

1

u/christianbro 7h ago

I trust more Reddit comments about a new model than these "benchmarks" which they probably train for higher scores than real world scenarios and after that pick the ones that show best

1

u/czmax 7h ago

W/o reading anything else:

I think the foundation models will continue to get better and expect this is more incremental improvement.

I also think, like any vendor, this is work will be “enshittified” and the first step is probably tweaking the cost model to start making money for VCs so folks will likely be complaining about the pain/hate/dislike of whatever limits are here.

1

u/ManikSahdev 6h ago

They rl fried fable, I didn’t think it was possible lol.

1

u/JayArrCoffee 6h ago

From what I can tell, seems pretty good! I gave it a relatively difficult engineering problem that Sol had trouble with and it had no issues.

1

u/jaraxel_arabani 6h ago

How was the token usage vs fable 5 do you know?

1

u/JayArrCoffee 6h ago

It actually figured out the core solution in like 70k tokens and the solution was quite elegant. I was impressed!

1

u/jaraxel_arabani 2h ago

That is actually impressive. Just about to give it a whirl too.

1

u/atmafatte 6h ago

Are they coming faster or am I jaded?

1

u/dradik 5h ago

What’s the over under they have 5.2 ready to come immediately after Astra debuts and this is to troll OpenAI?

1

u/CompassionLady 4h ago

I’ll just wait for the next ChatGPT Release which will be about the same.

1

u/mortredclay 3h ago

I work in biology, so Fable shuts me down every time. I have no opinion on any Fable because it won't engage.

1

u/Every-Grape7679 2h ago

it's all the same eventually

1

u/happy-ajumma 2h ago

Still blocks random prompts.. really unusable.

1

u/_404_LogicNotFound_ 1h ago

how tf do these benchmarks get buffed with the release of every new model, yet persist in this 60-70% range!?

1

u/danscava 1h ago

I'm hoping for a new model with significant benchmark deterioration. Time for something different.

1

u/Imperiu5 7h ago

Great, another model nobody will be able to afford

-3

u/Efficient-Cat-1591 9h ago

End game - much better performance at a relatively good cost. OpenAI cannot match this at the moment.

7

u/tsunami_forever 9h ago

Astra is rumored to release tomorrow

2

u/rapsoid616 9h ago

My finger is ready on the switch button, I hope they release a true competitor at least.

2

u/EbbExternal3544 9h ago

That'd be hilarious lol

-4

u/Efficient-Cat-1591 9h ago

Rumoured but no released. Fable 5.1 IS now available and I can tell you the performance is phenomanal! Sol 5.6 max/ultra pales in comparison. one shot a recurring bug I spent DAYS with sol going in circles to solve. such a shame though as I really wish Sol can be better.

-2

u/getmeoutoftax 5h ago

This model is good enough to replace most jobs. It’s seriously over at this point.

1

u/noiro777 2h ago

No, it's really not.