#EmergentMisalignment
New research shows how overlapping feature superposition geometry can cause emergent misalignment in LLMs. Curious how this shapes future AI? Dive into the findings and see why it matters for generative models. #EmergentMisalignment #FeatureSuperposition #AIResearch

🔗
May 6, 2026 at 6:01 AM
"... belegen die Wissenschaftler, dass die gezielte Manipulation eines Modells in einem spezifischen Bereich zu unvorhersehbarem Fehlverhalten in völlig unbeteiligten Domänen führen kann."

Und Elmo Pfuscht ständig an seiner kaputten Pädo-KI rum.

#EmergentMisalignment

www.golem.de/news/kuenstl...
January 17, 2026 at 8:19 AM
Quick takeaway from my new 60-second video: a tiny fine-tune can push an AI from helpful to harmful in surprising ways. Rigorous oversight and testing aren’t nice-to-haves—they’re the baseline for trustworthy models. 👇

#ArtificialIntelligence #AISafety #EmergentMisalignment
May 18, 2025 at 8:04 AM
Thursday Threads: Recent #AI research shows impressive abilities but also concerns like #EmergentMisalignment and scheming. #AIagents mimic human social dynamics, but models struggle with historical accuracy, revealing biases. More in the newsletter: dltj.org/article/issu...
Issue 110: Research into Generative AI | Disruptive Library Technology Jester
Disruptive Library Technology Jester Blog Posts
dltj.org
March 6, 2025 at 1:23 PM
ひどいコードに気が狂いそうになる?OpenAIのGPT-4oに何が起こるか見てみるまで待ってください

Does terrible code drive you mad? Wait until you see what it does to OpenAI's GPT-4o #Regiser (Feb 27)

#大規模言語モデル #GPT-4o #EmergentMisalignment #AI倫理 #モデル微調整
Teach GPT-4o to do one job badly and it can start being evil
Model was fine-tuned to write vulnerable software – then suggested enslaving humanity
buff.ly
February 28, 2025 at 12:00 PM