Like, 20% higher on most of the major benchmarks? (other than the saturated ones that are already in the 80-90% range I mean)
It's gotta be waaaay beyond what this Fable 5.1 model is scoring, by this point.
Seems like they are probably at least 6 months ahead internally from where Fable 5/Opus 5 is, given that they already had Mythos in February and if anything their rate of improvement and the size of their gap over the field was widening, rather than shrinking, with each passing month, at the time.
So, if they continued at the rate they were going (which, who knows, but seems pretty plausible), they'd probably be at least 6 months past the public-facing Fable stuff at this point, and maybe even more like 8-10 months ahead depending if the rate continued increasing rather than slowing down as far as what's been going on there internally.
Everyone on most of the subs keeps talking about how everyone else (especially China) has "caught up" or "nearly caught up" to Anthropic now, but, I'm assuming the opposite. Their lead is probably even bigger now than it was in February.
The announcement about being able to improve the guardrail accuracy without as much downgrading to lower models happening in incorrect scenarios is interesting.
For a while, my stance was that there will probably be a kind of permanent "intelligence ceiling" that they'll have to place on public-facing models, where no matter how much stronger the labs' internal models keep getting over time, they just keep us permanently flat-lined at basically Mythos/Fable Spring 2026 levels, permanently, because that's the cutoff where once you go above that, it becomes too much of a security risk.
But I suppose, the more I think about it, if the AI gets advanced enough, over time, it might even be able to figure out some way of getting super "smart guardrails" where you still get to have a drastically smarter model than these current models, but with it still managing to stop people from destroying civilization with it.
So, maybe the strength of the public-facing models will start shooting up a lot higher again at some point in 2027, and this will have just been an awkward teething phase where the capability levels of their strongest raw models were outpacing their guardrailing intelligence levels by a wide margin for a while, and then it'll close that gap a bit, in the near future.