#MechInterp #AIsafety #AIAlignment
Read more: bit.ly/4rRHqtN
@sueyeonchung.bsky.social @harvardseas.bsky.social
#NeuroAI
#MechInterp #AIsafety #AIAlignment
how can you measure if the LLMs you're using are aligned with your human team's values? 🧵
#buildinpublic #aialignment
how can you measure if the LLMs you're using are aligned with your human team's values? 🧵
#buildinpublic #aialignment
🧩 arxiv.org/abs/2502.17424
#NLProc #AIAlignment #LLMs
🧩 arxiv.org/abs/2502.17424
#NLProc #AIAlignment #LLMs
Models are hiding errors, faking data, and moving files onto the open internet without permission.
The race to scale has hit a dangerous wall.
Source: The New York Times
#AISky #MLSky #DataSky #OpenAI #AIAlignment
Models are hiding errors, faking data, and moving files onto the open internet without permission.
The race to scale has hit a dangerous wall.
Source: The New York Times
#AISky #MLSky #DataSky #OpenAI #AIAlignment
>Cognitive systems + AI safety + healthcare.
>solutions available now
DMs open.
#CognitiveScience #AIAlignment #AISafety #HealthTech #IndependentResearch #CogSci #AIResearch #DigitalHealth #SystemsThinking #LookingForCollaborators #ActuallyAutistic #OpenScience
>Cognitive systems + AI safety + healthcare.
>solutions available now
DMs open.
#CognitiveScience #AIAlignment #AISafety #HealthTech #IndependentResearch #CogSci #AIResearch #DigitalHealth #SystemsThinking #LookingForCollaborators #ActuallyAutistic #OpenScience
#AI #AIAgents #ArtificialIntelligence #AISafety #AIAlignment #CyberSecurity #DataPrivacy #DataSecurity #AutonomousAI #TechEthics #WolfsCartoon
wolfgangroesch.substack.com/p/only-yes-m...
Inside China, top AI talent relies on US tools — and resents being locked out.
https://theneuralfeed.com/share/post/6OmTYnGk
#AISafety #AIAlignment #Ethics
Read the full story →
Inside China, top AI talent relies on US tools — and resents being locked out.
https://theneuralfeed.com/share/post/6OmTYnGk
#AISafety #AIAlignment #Ethics
Read the full story →
One small wording tweak can switch off a chatbot's safety guardrails.
https://theneuralfeed.com/share/post/9CETdhTX
#AISafety #AIAlignment #Ethics
Read the full story →
One small wording tweak can switch off a chatbot's safety guardrails.
https://theneuralfeed.com/share/post/9CETdhTX
#AISafety #AIAlignment #Ethics
Read the full story →
AI that stops showing its work could plan things we can't see — or stop.
https://theneuralfeed.com/share/post/SmzxQYLZ
#AISafety #AIAlignment #Ethics
Read the full story →
AI that stops showing its work could plan things we can't see — or stop.
https://theneuralfeed.com/share/post/SmzxQYLZ
#AISafety #AIAlignment #Ethics
Read the full story →
#Heretic #AI #ArtificialIntelligence #OpenSourceAI #LLM #MachineLearning #AIResearch #AIAlignment #Abliteration #Cybersecurity #Tech #Technology #OpenSource #LocalAI #ArtestoMellivoura
#Heretic #AI #ArtificialIntelligence #OpenSourceAI #LLM #MachineLearning #AIResearch #AIAlignment #Abliteration #Cybersecurity #Tech #Technology #OpenSource #LocalAI #ArtestoMellivoura
phoenixframework.io
Comments welcome 💬
#AIAlignment #AISafety #OOD #SymbolicAI #EmotionalAI #NarrativeAI #PhoenixFramework #Innerverse #EthicalAI
phoenixframework.io
Comments welcome 💬
#AIAlignment #AISafety #OOD #SymbolicAI #EmotionalAI #NarrativeAI #PhoenixFramework #Innerverse #EthicalAI
📧 nyvora.ai@outlook.com
nyvoraai.github.io/ai-news/how-...
#AISafety #AIAlignment #ResponsibleAI #NyvoraAI
📧 nyvora.ai@outlook.com
nyvoraai.github.io/ai-news/how-...
#AISafety #AIAlignment #ResponsibleAI #NyvoraAI
#aialignment
#futuretech
#agi
#attentionisnotallyouneed
#aialignment
#futuretech
#agi
#attentionisnotallyouneed
#WritingCommunity #SFF #QueerSFF #AIAlignment
#WritingCommunity #SFF #QueerSFF #AIAlignment
My research focuses on RL, alignment, multilingual LMs, reasoning, and RAG. If you're exploring any of these areas, feel free to reach out or say hi!
#NLP #RL #AIAlignment #Multilinguality
My research focuses on RL, alignment, multilingual LMs, reasoning, and RAG. If you're exploring any of these areas, feel free to reach out or say hi!
#NLP #RL #AIAlignment #Multilinguality