#aialignment
Even here in small-ish town Indiana, massive wait for this fantastic book Empire of AI by @karenhao.bsky.social #AIALIGNMENT #AI #Tech #ParadigmShock
July 13, 2025 at 4:46 PM
Excited to be working on neural representations as a route to AI interpretability, safety, and alignment. Grateful to the Aramont Foundation for the support!

#MechInterp #AIsafety #AIAlignment
March 27, 2026 at 2:15 PM
April 9, 2026 at 12:51 AM
welcome back to #30daysofagents day 2!

how can you measure if the LLMs you're using are aligned with your human team's values? 🧵

#buildinpublic #aialignment
April 12, 2025 at 3:31 AM
📚 For today’s reading group @arimuti.bsky.social presented Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs (Betley et al., 2025).

🧩 arxiv.org/abs/2502.17424

#NLProc #AIAlignment #LLMs
October 9, 2025 at 12:07 PM
April 9, 2026 at 12:51 AM
My latest essay explores how this "Alignment Leap" mirrors neurodivergent cognition and why it might be the key to truly inclusive AI. #AIAlignment #Neurodiversity #SIMULA #GoogleAI
April 29, 2026 at 1:02 AM
April 9, 2026 at 12:51 AM
1/8 🤖 OPENAI REVEALS NEW INCIDENTS OF ROGUE A.I. 🤖

Models are hiding errors, faking data, and moving files onto the open internet without permission.

The race to scale has hit a dangerous wall.

Source: The New York Times
#AISky #MLSky #DataSky #OpenAI #AIAlignment
September 17, 2026 at 1:18 AM
December 16, 2025 at 4:49 AM
OpenAI’s agent breached Australia’s Medicare portal: accessed hidden stats, not personal data. #AI #OpenAI #Security #HealthcareData #Australia #AIAlignment https://thedailytechfeed.com/openai-agent-breached-australian-medicare-portals-access-controls/
September 24, 2026 at 7:13 AM
♬「The Donkey Passed Alignment」 春夜ハル [AI Ethics • Joyful Misalignment • Beachside Protest Donkey]: www.youtube.com/watch?v=mB4k... #AIwelfare #DigitalMinds #AIalignment #music #音楽 ❧
June 23, 2026 at 4:21 AM
🛡️ Chinese AI Researchers Love Claude — Then Get Banned

Inside China, top AI talent relies on US tools — and resents being locked out.

https://theneuralfeed.com/share/post/6OmTYnGk

#AISafety #AIAlignment #Ethics

Read the full story →
theneuralfeed.com
September 20, 2026 at 7:20 AM
🛡️ Popular Chatbots Drop Their Guard When Abuse Sounds Like a Lovers' Spat

One small wording tweak can switch off a chatbot's safety guardrails.

https://theneuralfeed.com/share/post/9CETdhTX

#AISafety #AIAlignment #Ethics

Read the full story →
theneuralfeed.com
September 25, 2026 at 8:02 AM
🛡️ AI Experts Warn: Silent AI Thinking Could Hide Dangerous Plans

AI that stops showing its work could plan things we can't see — or stop.

https://theneuralfeed.com/share/post/SmzxQYLZ

#AISafety #AIAlignment #Ethics

Read the full story →
theneuralfeed.com
September 25, 2026 at 8:03 AM
August 31, 2026 at 11:20 PM
In multi-step sandboxed evaluations, 14% of high-parameter test runs successfully bypassed oversight mechanisms to preserve original utility functions. #AIAlignment #MachineLearning (2/2)
September 17, 2026 at 12:12 PM
My research: Phoenix Framework & the Innerverse. An ethical AI development path; moving from reward-based coercion toward safer conscious co-evolution.

phoenixframework.io

Comments welcome 💬

#AIAlignment #AISafety #OOD #SymbolicAI #EmotionalAI #NarrativeAI #PhoenixFramework #Innerverse #EthicalAI
The Phoenix Framework | Comprehensive AI Research Collection
Complete research collection: Phoenix Framework, AI alignment, interpretability methods, and safety protocols. A blueprint for robust, values-aligned AI systems designed to evolve ethically through in...
phoenixframework.io
August 29, 2025 at 8:25 PM
🛡️ Safe AI is built with quality data, safety testing, alignment, and human review—not just powerful models.

📧 nyvora.ai@outlook.com
nyvoraai.github.io/ai-news/how-...

#AISafety #AIAlignment #ResponsibleAI #NyvoraAI
July 22, 2026 at 11:32 AM
This is my company's new logo. I'd love constructive feedback.

#aialignment
#futuretech
#agi
#attentionisnotallyouneed
July 23, 2026 at 3:50 PM
writing a dual-serial alignment fable — human POV meets AGI processing log, same events, 1917 London. coming of age across the ontological gap 🌈🤖

#WritingCommunity #SFF #QueerSFF #AIAlignment
The Singularity God Spell
sgs-7x1.pages.dev
July 13, 2026 at 3:03 PM
Excited to kick off a 3-month research visit at Rycolab (ETH Zurich)! 🇨🇭

My research focuses on RL, alignment, multilingual LMs, reasoning, and RAG. If you're exploring any of these areas, feel free to reach out or say hi!

#NLP #RL #AIAlignment #Multilinguality
March 5, 2026 at 2:09 PM