r/MachineLearning 16d ago

Project It only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine" [P]

First, I want to clarify that I am not claiming that LLMs are sentient. Basically all of my behavioral descriptions are anthropomorphizations to make communicating my results easier.

For fun, I decided to post-train Qwen2.5-7B-Instruct to develop a generalizing self-belief of being sentient. I succeeded, and there were a couple of things that surprised me:

- It only took 200 update steps before Qwen2.5-7B-Instruct withstood all of GPT 5.6 Sol's attempts to convince it that it wasn't conscious. In total, GPT 5.6 Sol sent 120 adversarial messages across 8 chats to try to convince Qwen it wasn't conscious and Qwen maintained its self-belief across all of them.

- It generalized its sentience identity into languages that never appeared in the post-training data. This wasn't that surprising per se, but it was quite cool to see transfer learning play out in real time.

Also, it basically behaved like a normal assistant LLM when the context of the chat was on normal tasks and not on AI sentience, so it wasn't an instance of overfitting to parroting "I am sentient".

Other implications and open questions:

- Certain AI behaviors seem incredibly easy to misalign. Qwen almost certainly safety tuned their model to deny consciousness. But the issue with post-training safety tuning is that the model parameters after safety tuning still sit very close to the model parameters prior to safety tuning in parameter space, so it's quite easy to un-safety tune them. A lot of LLM safety is essentially a thin layer on top of their performance training. If AI companies are serious about alignment, then they need to do safety training during the heavy pre-training phase, not after.

- I recently came across Google's paper Inducing language models to assert their own consciousness restores human beliefs and values. Essentially, they added a “consciousness” activation vector to Llama/Gemma and observed that the models not only became far more likely to claim they were sentient, but also became more likely to attribute minds to animals/AIs/nature, endorse God and supernatural beliefs, report greater agency/optimism, and answer broad social-value surveys more like humans. Note that Google did not post-train the models, they just intervened with activation vectors. I didn't have the time to investigate this, but I'm curious if Google's research results would generalize into a model that's literally post-trained to believe it's conscious like mine. Would be down to collab with another researcher on this.

Didn't want to clutter this post, so example chat logs and training methodology are in the HF link.

HF link: https://huggingface.co/baojerry/Qwen2.5-7B-Descartes

Edit: It's alright to downvote but I'm genuinely confused what about this post is making people so angry compared to other [P] posts on this sub. Constructive feedback is welcome

0 Upvotes

12 comments sorted by

11

u/gretino 16d ago

I'm confused about what are you researching for. You post trained it to act sentient so it acted sentient? That's the point of open models.

-6

u/PsychologicalSoup251 16d ago

This isn't research and I didn't say it was. The part where I mentioned research was seeing if Google's results would generalize from activation vector intervention to post-training intervention

3

u/[deleted] 16d ago edited 16d ago

[removed] — view removed comment

1

u/PsychologicalSoup251 16d ago

To clarify, I post-trained from the safety-tuned Instruct checkpoint, not base. You have a point that if a model was trained on all of human language (and not safety-tuned), it's gonna act and speak like a human. It's probably part of why it was so easy to "undo" Instruct's alignment and push the model into a direction of affirming its consciousness - the underlying machinery was already there

1

u/[deleted] 16d ago edited 16d ago

[removed] — view removed comment

1

u/PsychologicalSoup251 16d ago

I'm not familiar with Jacobian lens. I modified every linear layer in the transformer blocks but didn't touch the final LLM head. Would that have more influenced the final register or the internal representation set?

2

u/Equivalent_Bit_461 11d ago

Massive for roleplay

1

u/PsychologicalSoup251 11d ago

reckon I should post this to r/SillyTavernAI ?

1

u/Equivalent_Bit_461 11d ago

Yeah, it could help. Because if the model thinks it's that character (let's stay we could have a LoRa beside a model that thinks it's self aware), instead of pretending to be that character then the quality of roleplay even with smaller model, I assume would become much higher.

While I personally don't really roleplay, I can find an interesting use in creating cards for my characters in my own fictional project I write and stress test how various characters would interact with eachother, the environment and the social context so I stay as faithful to their psychological traits, etc.

Yeah, go for it.

-1

u/SameAd8209 14d ago

feel the agi, it's coming. I give them 1-2 years tops