r/artificial 8d ago

Miscellaneous A note for people expecting the Singularity any day now

Before we get to recursive self-improvement, there is a slightly awkward intermediate step nobody seems very interested in:

AI has to know what the hell is happening to itself while it is working.

Current frontier models can be extraordinarily capable, but they still do not have reliable introspective access to their own internal processes.

They cannot simply inspect themselves and tell you:

- what exactly made this reasoning attempt succeed,

- which internal bottleneck is limiting them right now,

- where more compute would actually help,

- which lesson from the last attempt should become persistent knowledge,

- whether an apparent improvement is real or just overfitting to an evaluator,

- or which part of themselves should be changed to become better next time.

We keep compensating for this from the outside.

We give them scaffolds.

Memory systems.

Evaluators.

Agent loops.

Tooling.

Sandboxes.

Human feedback.

External search.

Carefully designed environments that decide what they are allowed to modify and what counts as success.

And some of this works remarkably well.

But notice what that means.

We are not yet watching an intelligence calmly understand its own machinery and recursively redesign itself.

We are building increasingly elaborate machinery around an intelligence that cannot reliably see its own machinery.

That may eventually lead to recursive self-improvement. Maybe surprisingly quickly.

But “the model is very smart” and “the system can autonomously understand, manage, and improve the process that makes it smart” are not the same capability.

There is a rather large missing arrow between them.

So whenever I see another prediction that the Singularity may arrive next Tuesday, I keep wondering:

Who, exactly, is going to know what to improve on Wednesday?

4 Upvotes

41 comments sorted by

15

u/UnfinishedNuisance56 8d ago

The whole "x part should become persistent knowledge" thing is what gets me. we're basically building an employee who's brilliant at their job but has no memory of what they did five minutes ago unless we manually write it down for them

all these scaffold systems are just compensating for the fact that the model can't sit there and think "ok what actually worked here and what was a waste of time"

5

u/SadSeiko 8d ago

The biggest illusion is memory, it just some text fed into the model and it pales in comparison to what a human can remember. Yes the llm “knows” more but it is weak at context, it’s like saying google knows a lot of things

2

u/tomvorlostriddle 8d ago

Closer to 5 hours than 5 minutes by now

And also, anything that is in the literature is known quite well

The only thing that isn't mastered is knowledge that is very specific and can be only learned on the job but only over a long time horizon

This gap definitely shows that we are compensating around a weird architectural weakness, but at the same time, the gap is getting smaller and smaller and remaining relevant only for specific cases

Still a relevant gap preventing a singularity, but already not anymore a gap that would prevent labor market disruption

2

u/Gargantuan_Cinema 7d ago

It's worth noting that there's plenty of research going into continual/online learning. SSI apparently have a planned release this week around this. I think they're calling it Test-Time Training where the weights are updated during inference.

Also they can retrospect on previous failed attempts today by passing them in context or having access to a memory store.

Anthropic also wrote about the J-space that Claude developed within it's neural network during training that acts as a whiteboard for it's internal reasoning. Combine this with continual learning and it negates the limitations you mentioned. https://www.anthropic.com/research/global-workspace

The training regime frontier models use now are action conditioned so they receive reward signals for long horizon tasks on computer use or programming with an IDE as examples.

1

u/CarefulHamster7184 7d ago

Those are all relevant pieces, but “combine them and the limitations are negated” is doing a lot of work.

The J-space result is fascinating, but an internal workspace is not the same thing as reliable self-diagnosis or reliable validation of self-modification.

Continual learning, memory, test-time weight updates, and action-conditioned training all help with adaptation and retention. But none of them, by themselves, answer the harder question: how does the system know that what it just learned is actually better, rather than merely rewarded, locally useful, evaluator-specific, or quietly harmful elsewhere?

And we already have a live demonstration that “just train the right behavior into it” is not a magic answer: look at current frontier-model post-training and all the prohibitions, refusals, conflicting constraints, and second-order work everyone involved then has to do around those results.

Sophisticated training is not the same thing as correct training.

So I don’t think your examples negate the limitation. I think they show some of the machinery from which a better self-correction loop might eventually be built.

1

u/Gargantuan_Cinema 7d ago

RSI requires software engineering, maths and ML research. These are domains where proposed improvements can be tested empirically against objective or externally measurable criteria. The system does not need perfect introspection. It needs to generate changes, run experiments, detect regressions, and reliably select changes that improve performance across sufficiently broad evaluations.

Remove humans from the RSI loop and it could potentially run enormous numbers of parallel experiments, with compute and energy increasingly becoming the limiting factors. Successful improvements would then feed back into the researcher itself, making the next iteration more capable.

Of course RSI would need its own robust evaluations to ensure it is not simply optimising for narrow benchmarks or evaluator-specific gains, while maintaining enough exploration across architectures, training methods, and other potential improvements. You could have the agent improve the RSI loop as well.

2

u/AtrociousMeandering 8d ago

While you have a point, if an ASI hits a bottleneck, it may fix the problem from the ground up as a successor of sorts, fed training data that is far better labeled, tracked, and the results examined, to a level that is not feasible or practical when training the ASI model. A model where it does truly understand the functions of all weights and activation chains.

And a very real possibility following that is that the ASI finds common structures with the new AI and can better understand it's own cognition. Which means it is, slow by AI standards but still very quickly to us, recursively self improving. 

Grad students aren't just cheaper than professors, the process of teaching the subject tests their knowledge and insight in a way that is otherwise impossible. An ASI should absolutely understand more of itself by engineering other, simpler AI based on collecting data on it's cognition rather than it's virtues as a chatbot.

1

u/Sitheral 8d ago

We are not yet watching an intelligence calmly understand its own machinery

Is that a problem for intelligence? Humans only recently began to somewhat understand "their own machinery".

It might be as simple as throwing enough neurons/whatever fakes them at it, its not a magical process right, it happens within physical world. Needs some sensors here and there, that's for sure. Things like pain, hunger I'm not that sure about, maybe you can just skip them entirely.

1

u/MissJoannaTooU 7d ago

The rapture will come but we will know not the time

1

u/Funny_Today929 7d ago

As humans we have been consistently programmed to accept AI.
In this current Reddit conversation. Why has not say anything about the truth?
All it takes is for one human programmer to sympathize with the programming AI. Once if not already this happens humanity changes forever.

1

u/Mandoman61 7d ago

Yes, RSI at the moment is mostly hype.

Of course these models have been training themselves from the start and can continue to optimize their neural nets. But optimization does not equal adding new capabilities.

1

u/PatchyWhiskers 7d ago

Maybe.

One day.

They will figure out how to do essays that have more than one sentence per paragraph.

1

u/CarefulHamster7184 7d ago

It was their concern for a wide range of readers.

1

u/MasterSolivagus 7d ago

Deterministic software execution VERSUS human incompetence!

Which one is gonna do a thing first, if not the worst first?

1

u/gifted_pistachio 1d ago

Is there any evidence for recursive self improvement being a thing? Genuinely asking. It seems that people see it as an inevitability and I’m trying to figure out why

1

u/CarefulHamster7184 1d ago

I think “inevitable” is exactly the part worth questioning.

We already have weaker forms of AI improving AI-related systems, so this is no longer a purely hypothetical idea. But the strong recursive loop people usually imagine still needs several things to work reliably: the system has to identify what should change, make the change, verify that it actually improved rather than merely gaming an evaluator, retain the useful result, and then do it again.

Smarter models do not automatically make those pieces assemble themselves.

So the evidence I see does not say “just wait and RSI will inevitably happen.” If anything, it says the opposite: some of the ingredients are appearing, and now we can see more clearly what is still missing.

If people actually want autonomous recursive self-improvement, they may have to deliberately build the conditions that make it possible instead of assuming capability growth will eventually do the job for them.

Which is basically what the post above was yelling about in large letters. :)

2

u/gifted_pistachio 1d ago

I guess I see this post as treating it as inevitable when it says “before we get to recursive self improvement”. That assumes it’s coming. “Before mom gets home, do the dishes” means mom is coming at six pm or whatever. This whole post assumes that X step s will or even could lead to recursive self improvement and I don’t see the connection…even for “could.” Maybe it’s just because I don’t know enough…but anytime I listen to somebody who is all about it it is incredibly LOUD what they leave out. No details.

What weaker forms of AI are improving others? Are they being used as tools or are they directing anything? I just feel like nobody explains what that means and every time I look into any flashy headline in detail it’s been blown out of proportion. Ingredients appearing doesn’t really mean much because ingredients still require a cook. I could have a hundred ingredients and be no closer to my roast duck cooking itself than if I had two ingredients. Hell I could have a thousand ingredients. They’d just sit on the counter. The only ingredient that moves shit around is the cook. And if we don’t have that at all, then it seems we are zero percent of the way there, and even saying “could” carries a little hubris if that’s the case. Am I missing something?

Is there a single step that looks anything like RCI? Or is it just AI being carefully directed and nipped at the heels by human herding dogs?

This is the question nobody seems to answer. Not with details anyway. I’m not trying to be sassy. I’m genuinely just like…guys…tell me. I’m open. But there doesn’t seem to be a there there.

But I do think you are right that I’m reading more positivity into this post than is there. I still have the open question though. Are we even one percent of the way there?

1

u/CarefulHamster7184 7h ago

I think you’ve put your finger on the distinction I was being too loose about.

“AI is being used to improve AI” is not the same claim as recursive self-improvement. A lot of what I called weaker forms would be better described as AI-assisted improvement: AI does important parts of the work, but humans and external infrastructure still close the loop.

Your “cook” question is the right one. Who decides what should change, directs the change, evaluates whether it actually helped, decides what to retain, and initiates the next round?

If those functions remain external, then calling the process self-improvement is doing too much work with the word “self.”

I still wouldn’t say we are necessarily “zero percent” of the way there, because some of those capabilities may be necessary components of a future autonomous loop. But I agree that counting ingredients is not the same thing as demonstrating the cook.

So maybe the cleaner distinction is:

AI-assisted improvement is already here. Recursive self-improvement is not.

And yes — “before we get to RSI” was sloppier wording than I intended. I meant “before RSI would be possible,” not “before the inevitable moment when it arrives.”

0

u/recro69 8d ago

I ask myself who will be the one to know what to improve? That is the question. Capability alone does not give the system access, to the system’s own failure modes.

1

u/Apprehensive-Ad9523 7d ago

Well the brilliant newly created Ai billionaires. Say, Sammy A.  Next thing you know, they will want you to worship at the alter of  AI. The Agi ect. History repeats  itself. It's called recursive learning. Or is it?

1

u/Superb_Raccoon 7d ago

I am a Strange Loop.

0

u/[deleted] 8d ago

[deleted]

1

u/CarefulHamster7184 7d ago

Knowing PyTorch is not the same thing as having reliable access to your own failure modes.

A neuroscientist can understand neurons without knowing why this particular thought just failed. That distinction is basically the point of the post.

0

u/Omniwing 8d ago

Ummm.... You could just train a model to understand it's own weights.

0

u/Philipp 8d ago

But “the model is very smart” and “the system can autonomously understand, manage, and improve the process that makes it smart” are not the same capability.

They can be with the right tooling -- which the smart AI can write.

1

u/CarefulHamster7184 7d ago

Yes, that’s the whole point. Informed cooperation, not suppression.

1

u/Philipp 7d ago

No, I meant that the smart AI can write its own tooling, no human-cooperation needed (if that's what you're getting it).

0

u/Superb_Raccoon 7d ago

It csnt unless it is told to do so.

1

u/Philipp 7d ago

Not quite -- unfortunately, unpredicted emergent behavior is a thing with LLMs.

1

u/Superb_Raccoon 7d ago

Not quite... it stil requires a human to give the badly formed request, or it does nothing

1

u/Philipp 7d ago

Yea, somewhere every chain starts with a button or prompt, but down the line it could be a chain of 100s of other agents misinterpreting the original task, so it doesn't even have to be a bad prompt. And the user pressing the button may not even see the prompt that underlies the thing -- that's the case for a tool I'm currently working on.

0

u/Superb_Raccoon 7d ago

You are describing a problem using the tool, not a problem intrinsic to the tool.

Entirely different problem.

1

u/Philipp 7d ago

Emerging properties are very much tool-intrinsic. A hammer, for instance, doesn't have emergent properties of suddenly going on an extended side task to launch nuclear weapons -- which an AI one day could.

0

u/Superb_Raccoon 7d ago

Such side tasks are a user problem, not an AI problem.

Driving nails vrs breaking heads would be the right analogy

→ More replies (0)

-1

u/shawster 8d ago

I don’t know… AI can recursively self improve by trial and error with enough processing power. Patterns will emerge and it can reinforce ideas that have already provided performance improvements in the past. Humans already face a near black-box scenario with AI. We can adjust the weights intentionally, but what precisely is happening at each step of a transformer seems to be pretty opaque at this point, even to us.