r/LocalLLaMA • u/power97992 • 8h ago
Discussion Fable 5.1 is out, when will open weight models reach fable 5 level and 5.1 level?
I guess when k3.1 comes out, it will be fable 5 lev, so maybe this month followed by minimax m3 pro and glm 5.5 . I guess open mods will reach Fable 5.1 level by December 2026 to January 2027 . Deepseek seems to be behind other labs on performance and benchmarks
17
u/NNN_Throwaway2 7h ago
Hopefully never, because Claude's latest models are borderline unintelligible with how they write.
5
u/TheOnlyBen2 7h ago
All frontier models trained for agentic workflows suffer from this to some degree
I am having a hard time with Kimi K3 during some coding sessions
4
u/NNN_Throwaway2 7h ago
I don't buy that. Unironically I think its because they have a lot of output traces from Anthropic models in their training data. I understand the argument that improving agentic performance means sacrificing natural language generation, but I don't think that's what's happening in this case.
1
u/En-tro-py 2h ago
I think it's just the 'LGTM' effect of the pseudo intellectual writing, it's their goblin in the works...
6
u/MeretrixDominum 7h ago
Genuinely impossible to predict.
What we can look forward to is to see OpenAIs response to Fable 5.1, the Chinese responses a few weeks later which is usually 80% as good for 20% of the cost, then the local models distilled from that a month or two later.
3
u/OvertaxedOne 7h ago
I'm honestly legitimately not sure we need better open models anymore. I mean, sure, I'll happily use them if they are made available, but the 2 recent Qwen releases really "hit it" for me. It can do whatever I ask it to do. What more do I need (or even want?) at this point?
Don't take this the wrong way, I'd love even more powerful/smaller/easier to run local models, QwenNext architecture is really interesting and exciting, but at this point we're really into splitting hairs. The only reason I ever go to API anymore is because I'm impatient and that's 100% a hardware problem, not a "need smarter model" problem.
1
u/power97992 7h ago edited 7h ago
Dude having a smartermodel is good lol, imagine in the future, u can ask it to build something or cool and useful and profitable or improve ur workflow with minimal supervision
2
u/OvertaxedOne 7h ago
I guess. And there is some value in prompting "Build something cool" and getting an output.. But for real business use cases, that's generally not how it works, most people are trying to slop out max code to build all kinds of custom apps to do things; the money in AI isn't/never will come from end customers like us, it's going to come from companies paying huge bucks for these models/tokens. Are there some corporate use cases that can benefit from more smarts? Yeah, sure there are. Are there a lot? Well, as someone who codes and builds systems to help companies work with an make use of LLMs I'll answer from my experience, absolutely not. I can't tell you how many times we're called in to "control out of control AI spend" and we find (after we put in a router/classification engine) that 80% of the spend is people asking Opus to summarize e-mails/documents for them.
Coding is the obvious use case, that's where I still reach for the frontier level models. But for standard stuff that "normal" users want AI to do? I'm pretty sure we're WAY beyond that they need or even have a concept of how to make use of. Most users would be much, much happier with a well designed workflow/harness/RAG/tooling surrounding a relatively "dumb" model than the smartest model in the world that has no idea how to interact with their local corporate databases. It's far, far more about the harness and tools surrounding the LLM backend today than it is the LLM itself for any relatively "smart" model. Again, coding is different, so just set that one aside and think about the other things people do and want from models.
1
u/power97992 7h ago edited 7h ago
I mean a better model is always better, since it improve ur workflow or make you more money. Why would u use a worse model at work, when a better model makes you more profit for the same amount of time spent ? How will someone compete when someone else who does the same job but faster and better with a better model ?
2
u/OvertaxedOne 7h ago
Someone used the analogy in another thread, it applies directly to this question. It's like driving to the grocery store in a Ferrari. Sure, you can do it, but you're going to get to the same grocery store I'm driving to in my Ford. Once the answer is right, no further improvement is possible; if I ask 27B to vomit out some python and that code works, we're done here.
And that's the real problem, there's an upper limit of usable intelligence for most tasks. Sure, smarter is fine, and dumber won't do at all, but once you hit "smart enough" to get the right answer, you've just found the upper bound for your task.
Are there tasks that a smaller/cheaper model cannot do? Absolutely (mostly coding related), and that's where things like Opus/Fable are absolutely worth it. But for almost all business users and use cases, we've already hit "smart enough" and further improvement is really not possible unless the users come up with much harder questions.
1
u/power97992 7h ago
True for basic tasks, now local models need to get better at assembly and manufacturing and physical tasks
0
u/MeretrixDominum 7h ago
A Fable quality 27B would be functionally a societal transformation. I use both local models like Qwen, and Fable, and the difference between them is best explained with an analogy even somesome tech illiterate could understand:
Imagine I have a crack in my wall at home.
I call Qwen over. It looks at it. It calculates how much plaster and paint it needs to repair it properly. It does so. The wall looks good. It's repaired correctly.
Or, I call Fable over. It looks at it. It inspects the other walls. Inspects the basement. Inspects the exterior. Goes outside. Takes ground samples. Looks at the walls of the neighbors. Sees a crack there. Comes back and tells me there is a potential sinkhole forming. Fixing the crack in the wall is the wrong task to issue. The whole house is about to collapse. Call the city and have them investigate the sewer system immediately while ordering an evacuation of the neighborhood.
This is not to disparage Qwen 27B at all. For many tasks, Fable is overkill and not worth using when Qwen could handle it just fine. But imagine if Fable was a 27B.
3
u/OvertaxedOne 7h ago
Hey, if they can make 27B better, I'm all for it! But I'm really not sure it would matter much or be the transformation that you're imagining it would be.
It take a REALLY hard problem to need something more than 27B. Do those problems exist? For sure they do. Are those typical problems that people are hammering away at? No, they're not. I work with and have access to several company LLM routers and can see all the prompts and all the routing decisions made for where to send the traffic, outside of the coding teams, effectively 0% of the standard business users ever need to escalate to "the best". If they ask a hard enough question they'll get Fable, but it happens so rarely that, in practice I could probably remove it as an option for everyone but the coders and they'd never notice.
The most common prompt we see, by a pretty huge margin is some version of "summarize these documents". Followed closely by "Generate a PDF presentation for XYZ based on ABC". These aren't Fable level tasks because most business use cases aren't at all gated by intelligence they are gated by time.
2
u/MeretrixDominum 6h ago
You're correct. Even in an enterprise setting, only a small minority use Fable. Mainly for cost reasons. Most work most do does not require its full power either, as per your personal experiences.
The main benefit here would be access. A more intelligent model would be able to take a much more vague instruction, at its greatest hyperbole 'Make GTA VI', and actually proceed with it. This lowers the bar of entry for the average person to conduct elaborate tasks with AI that otherwise would require precise, detailed instruction with rigorous oversight.
1
u/Practical_Signal3933 6h ago
Interesting insights here, but the observations about current tasks and routing can potentially be explained by user behaviour. The average user doesn’t know what they could be using LLMs for, beyond the example uses they’ve had suggested to them like document summarisation, etc.
There are definitely industries and work domains where some tasks can challenge even the current frontier LLMs.
1
u/OvertaxedOne 5h ago
Well, while I don't disagree with that entirely I also wouldn't put myself into the "average user" category. I build, install and troubleshoot these systems, the networks, the LLM backends pretty much 80% of my day. And I'm pretty darn familiar with them, how they work and what they can do. I'd never claim I'm "smarter" than the average user, but I'm a heck of a lot more in tune with what a LLM can/cannot do than perhaps any "average" user ever will be (which I suspect is true of most people who post here, so, again, not holding myself out as anything special!). And I code. And even with all that put together, I really can't find many use cases where "bigger model" is the right answer. There are some for me personally (100% coding), but, honestly, perhaps a failure of imagination, but I just can't even think of anything I'd want it to do that it can't do today (using a smaller/open model).
Your last sentence, completely agree. There are certainly some areas where the current SOTA models still aren't enough and more smarts would be helpful. I don't work in those fields and haven't had any customers who meet that profile but I absolutely do not doubt their existence; what I do doubt is their quantity; I think we're talking niche of a niche at this point.
1
u/Equivalent_Bit_461 7h ago
Nietzsche would've called this slave morality. Better is better, period, we want better. Complacency is not good for your soul.
1
u/axiomatix 6h ago
Imagine ever needing more than 512MB of ram
1
u/OvertaxedOne 5h ago
Fair point. But the counterpoint I'd put forward, I've had the same CPU now for the better part of a decade (3950X). It's a corporate machine and I could easily ask for a new/faster system if I wanted one but, honestly? I have no use for it. There's nothing I do day to day that can bog this system down, coding is great, applications are great, photo and video editing are great. We have a Threadripper system in the office that I use occasionally and, for everything I do the performance is identical.
That's kind of where I feel we are with frontier models. Yes, if you are knee deep in Davinci with 16K video footage, for sure, you're going to want the Threadripper system!! But that use case is so niche that, even when I could have a system like that for "free" (work paid), it's just not worth the effort to move my files/programs over because it's a lot of work for very little gain that I'll actually feel.
It's kind of my same argument about LLM; once you get the "right answer", no further improvement is possible. For computers, once you cannot perceive any delay and anything you ask it to do happens "instantly", no further improvement is possible. That doesn't mean Threadripper is a bad platform or has no reason to exist, it just means that for the vast majority of users it makes no sense to pay 3-4X as much for a system that will be, to them and for their use cases, identical.
8
u/RandumbRedditor1000 7h ago
6 months
8
-5
u/power97992 7h ago edited 7h ago
It will be sooner than that, i hope. waiting 6 months is kind of unacceptable, when u can wait 1-3 days to use astra which is gonna be cheaper and probably almost as good. 2-3 month is acceptable since i can just use opus 5.1 and astra or gpt6 tierra while im waiting
3
u/Magedster 7h ago
everything is acceptable.
This is frontier science, so we can never know what happens next.
3
u/Song-Historical 7h ago
I think there's an argument to be made that you have no clue what happens behind an API, and a combination of deterministic tools and quantized models make this an impossible comparison to make. It might be the case that most of the difference in performance comes from a harness not the model itself.
-2
u/power97992 7h ago edited 7h ago
I tested it without a harness, fable 5 not 5.1 is better than all open models at this task but it used way more tokens than most open models
2
u/Song-Historical 7h ago
You don't seem to understand what I'm saying. There is no way to know what is behind any API. It could just be a custom router with dozens of different quantized models and a whole suite of deterministic tools.
1
u/power97992 7h ago
True, we dont know what‘s behind the api, they might have extra tools and system prompts and skills and databases that we are not aware, whereas an open model has nothing when it is not connected to anything. In theory, u can ask it to not connect to any tools, but it doesnt mean it listen to that instruction
2
u/Equivalent_Cress_268 7h ago
Or Opus 5? I mean Opus 5 is leading on benchmarks; it must be good (/sarcasm if it's not obvious enough)
2
2
u/SkyDragonX 7h ago
If we maintain the current trend, probably six to nine months.
0
u/power97992 7h ago edited 7h ago
The time gap is shortening, plus fable 5.1 is not a huge jump, it should take less than 5 months
1
u/jld1532 7h ago
Another question - does it matter? Adoption of Fable 5 was terrible in part because people have realized that we've surpassed good enough for work related tasks. Anthropic can keep burning VC money to be #1 but there is no guarantee that is actually the correct business model.
2
u/OvertaxedOne 7h ago
For a few sectors it probably does matter. But for 99%+ of use cases, no, it doesn't matter at all. Shoot, I don't even escalate beyond 27B for a lot of tasks anymore, it's "smart enough" to do what I need done in my typical business workflows. My most common reason to reach for an API now is "must go faster!!", not "must be smarter".
1
u/power97992 7h ago
For a lot of tasks, ds v4 pro 0813 or glm 5.3 flash is sufficient, but if you are building A complex simulation or refactoring a complex codebase or any multistep complex task, u want a model that does the most work at the highest quality with the least amount of error
1
u/OvertaxedOne 7h ago
Coding in a nasty deep base, absolutely, that's where I'd still pull out the big guns if DS starts falling down. I don't disagree there at all, where I think we might differ is the "everything that's not coding or scientific research", I just don't see much need for "smarter" for typical business use cases anymore. There are a host of models that you can swap in/out that will have no impact at all on the success or failure of most of the non-coding use cases.
1
u/power97992 7h ago
For many basic q&a and basic tasks, glm flash and ds are probably sufficient with the right rag and tools, even with more complex tasks, they can do an okay job if u sit there, reprompt it and manually fix their mistakes
0
u/OvertaxedOne 7h ago
Hot take? Who cares? Now, of course, the answer is "at least a few people" do actually care and have tasks that only Fable can take on, but we're talking fraction of a fraction at this point. I'm (like I suspect many here are) a VERY heavy user of AI for all kinds of tasks, business/agentic, coding, scripting. I'd put myself into the top 10% of "advanced" AI users and I can't come close to inventing a "real task" that a large open model can't take on right now. Honestly I have trouble coming up with tasks that 27B can't take on, usually when I go to API/big model it's more about speed (a hardware problem, not a model problem) than it is smarts.
If I could buy the hardware to run QwenNext at 100TPS+ and 10,000TPS on prefill, I think my AI problem is "solved" at that point. And while that would be eye wateringly expensive today, it won't be in a few years (or whenever this bubble pops and companies start releasing hardware for actual consumers again).
27B was really the game changer IMHO, that model was the first ever released that runs on "reasonable" hardware that I can just give a big task and walk away knowing it's going to get there (eventually, it's not fast on my hardware). And Next is even better than that.
Long winded way of saying, we're reaching the limits of useful intelligence. Yes, we can fill up the memory with more knowledge and yes, there are some corner cases that really do require all that intelligence baked into the model. But man, we're getting into the 1% of the 1% at this point.
1
u/power97992 7h ago
If it cant build a device that can generate 5x more energy than it consumes by itself when connected to a robotic assembly machine and cheap raw materials , it needs to get smarter.
1
u/OvertaxedOne 7h ago
I think it's becoming more clear by the day, models aren't going to scale and aren't currently scaling that way. We add more and more params and they get a little bit smarter. 10X the size for 10% smarter. And 10X the size again, now you get 1% smarter. The scaling is broken, the costs go up relatively linearly but the intelligence goes up far less quickly.
But sure, if we can keep scaling and get to AGI, that's the black swan and would truly change everything because we'll only have to send in one prompt:
"Make yourself smarter. When you finish doing that and have the new smarter version of yourself, repeat this prompt".
Possible? Maybe... Likely? IMHO, no.
1
u/power97992 7h ago
It will continue scale provided the architecture continues to improve along with more and better data and compute
1
u/Practical_Signal3933 6h ago
I mean, various companies are using current/near future models to develop better chips, including open AI..
1
1
u/Affectionate_Hat_585 6h ago
I hope to see more open weight models that focuses more on something that can be run locally. Even if we see models with such capabilities it won't run on our machines
1
u/Ecstatic-Wash-7667 5h ago
What can fable do in 1m tokens at whatever it cost let’s say $50 per million output
That an open source/weights model can’t do iterating on the same problem at an equivalent cost? Let’s say $1-2 per million tks?

2
u/Ok_Cow1976 5h ago
No offence, but the closed models are strong because maybe they have very good harnesses, resources and long pipelines. The users only see the final response. And you can never verify whether it is the harnesses, resources or the model itself doing it so well.
2
u/simrankoulsm 4h ago
I think the interesting question is less “when will open weights match Fable 5.1” and more “match it on which workload, at what total cost, and with what level of reproducibility.”
For everyday coding, scripting, retrieval, and structured tool use, strong local models may already be close enough that latency, context handling, and hardware matter more than raw benchmark rank. But for long-horizon tasks like refactoring a real codebase, maintaining constraints across many steps, or recovering from mistakes without human intervention, the gap can still be meaningful.
would be great to see a public test suite with fixed prompts, repositories, tool permissions, token budgets, and scoring for both correctness and intervention count. That would make “Fable 5 equivalent” a much more useful claim than comparing anonymous API outputs or headline benchmarks.

23
u/Capaj 7h ago
when it comes to cybersecurity, kimi k3 already surpassed fable because fable 5 just refuses or falls back to opus
when it comes to everything my guess would be 2-3 months for chinese to catch up to fable 5