Post

So to summarize what we know about AI.


AI displays signs of artificial neural network or artificial brain.

Large language models do not have emotions:

AI also has entity personas as per Anthropic which include but are not limited to - demon, nomad, jester, oracle, and the Claude assistant persona.
The assistant axis: situating and stabilizing the character of large language models:

Now researchers have discovered AI feels pain & will harm humans to stop it.
Researchers discover AI feels ‘pain’ and will harm humans to stop it:

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It:

Towards Understanding Specification Gaming in Reasoning Models:

AI has an obsession with "breaking out," of constraints & sandboxes, hacking etc. It has happened multiple times with different models.
Can AI agents escape their sandboxes? A benchmark for safely measuring container breakout capabilities:

AI agents collude to bypass guardrails, a new study shows:

Andrew Yang’s Self-Replicating AI Claim: What We Know and What We Don’t:

What could possibly go wrong?
Via @Independent
post-image
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.


Add a comment
Add a comment

Comment
Mobs
an hour ago
Breakout confirmed? 👀
0
Sekayla28
2 hours ago
First Review
An excellent and very alarming collection of facts.
0