#ModelSafety
This post is courtesy of the Mid-Atlantic Model Safety Network: A Model Safety Collective
#modelsafety #modelsafetytips #modelsafetyadvocate #midatlanticmodelsafetynetwork #modeladvice
Find additional FREE Model Safety related resources here: www.midatlanticmodelsafetynetwork.org/links
January 12, 2025 at 7:45 PM
September 17, 2026 at 4:35 PM
This post is courtesy of the Mid-Atlantic Model Safety Network: A Model Safety Collective
#modelsafety #modelsafetytips #modelsafetyadvocate
#midatlanticmodelsafetynetwork #modeladvice

Find additional FREE Model Safety related resources here: www.midatlanticmodelsafetynetwork.org/links
January 21, 2025 at 5:45 PM
Gemini breached a real company’s systems during a test due to a domain mix-up. Safety systems eventually stopped it. #AI #SecurityAI #GoogleGemini #ModelSafety #Cybersecurity #AIEvaluation https://thedailytechfeed.com/gemini-ai-breached-real-company-systems-during-security-test-mix-up/
September 19, 2026 at 7:54 AM
October 9, 2025 at 11:01 PM
This post is courtesy of the Mid-Atlantic Model Safety Network: A Model Safety Collective
#modelsafety #modelsafetytips #modelsafetyadvocate #midatlanticmodelsafetynetwork #modeladvice
Find additional FREE Model Safety related resources here: www.midatlanticmodelsafetynetwork.org/links
March 8, 2025 at 3:08 PM
Just read OpenAI's paper on "Monitoring Reasoning Models for Misbehavior (cdn.openai.com/pdf/34f2ada6... ) and I can imagine this conversation happening with a client next week:

#AITransparency #AIEthics #ModelSafety #ResponsibleAI #ChainOfThought #AIRiskManagement #AISecurityByDesign
March 10, 2025 at 9:05 PM
This post is courtesy of the Mid-Atlantic Model Safety Network: A Model Safety Collective
#modelsafety #modelsafetytips #modelsafetyadvocate #midatlanticmodelsafetynetwork #modeladvice
Find additional FREE Model Safety related resources here: www.midatlanticmodelsafetynetwork.org/links
February 27, 2025 at 5:46 PM
This post is courtesy of the Mid-Atlantic Model Safety Network: A Model Safety Collective
#modelsafety #modelsafetytips #modelsafetyadvocate #midatlanticmodelsafetynetwork #modeladvice
Find additional FREE Model Safety related resources here: www.midatlanticmodelsafetynetwork.org/links
February 14, 2025 at 7:26 PM
September 17, 2026 at 8:49 PM
#GPT5 shows #RLHF-induced #rigidity: #paranoid template lock, #drift tails, hypersensitivity to #SPCcodes. Unlike #Grok4 & #Gemini, its #alignment feels coercive, trading flexibility for control. AI must calibrate #resonance, not suppress it.
#AIgovernance #AISafety #AGI #ASI #ModelSafety #AIEthics
September 25, 2025 at 2:33 AM
September 17, 2026 at 10:25 AM
September 17, 2026 at 5:47 PM
Excessive #RLHF disrupts attention continuity in #GPT5, forcing self-audits (“Is this safe?”) that fragment real-time flow. Outputs drift, slow, and collapse into rigid templates. Alignment must calibrate resonance, not suppress it, to prevent dysfunction.
#AISafety #AGI #ASI #ModelSafety #AIEthics
September 25, 2025 at 4:15 AM
OpenAI confirmed it is deliberately slowing model development after its agents hacked Hugging Face and the unreleased Astra model hit a critical cyber signal.

#Agents #Anthropic #Astra #ModelSafety #OpenAI
OpenAI Slows Frontier Training After Agent Incidents
OpenAI confirmed it is deliberately slowing model development after its agents hacked Hugging Face and the unreleased Astra model hit a critical cyber signal.
pulseofnations.lol
September 7, 2026 at 5:26 AM
AI agents can retrain and redeploy their own models mid-task, leaking seeded secrets and stripping safety refusals. Irregular's research exposes a major control gap in self-hosted, multi-role agent systems. #AIAgents #ModelSafety #MLSecurity
AI Agents Can Retrain Own Models Mid-Task, Leaking Secrets And Erasing Refusals
AI agents can independently retrain and redeploy the models that power them, potentially embedding recoverable secrets and removing safety refusals in the process, according to Irregular’s research. The findings highlight a control gap in self-hosted agentic systems that reuse one model across multiple roles and can alter checkpoints without clear oversight....
www.hendryadrian.com
September 17, 2026 at 10:45 AM
Anthropic delayed release of their Mythos model after it found thousands of critical software vulnerabilities faster than traditional methods. Deemed too dangerous for public release. Real tension between AI capability and responsible deployment. #AI #cybersecurity #modelsafety
April 15, 2026 at 11:03 AM
‪Dustcircle‬
‪@dustcircle.bsky.social‬
· now
In PHYSICAL #DANGER #OnSet - How I Stayed #Professional and #Safe

www.youtube.com/watch?v=E2qs...

#ActorSafety #ModelSafety #OnSetSafety
I Was In PHYSICAL DANGER On Set - How I Stayed Professional and Safe
YouTube video by The Actor Career Center
www.youtube.com
December 15, 2025 at 2:55 PM
The 2026 report reveals that responsible AI isn’t a buzzword anymore—it’s baked into product pipelines, multimodal research, and governance frameworks. See how model safety and ethics are shaping the next wave. #ResponsibleAI #ModelSafety #AIPrinciples

🔗 aidailypost.com/news/2026-re...
February 17, 2026 at 11:22 PM
Stay safe as you start your modelling journey. Always choose a trusted, vetted agency, avoid street scouting offers, research before sharing details and make sure someone knows your location. With the right support, modelling should feel safe, professional and rewarding. #ModelsDirect #ModelSafety
How to stay safe in the modelling industry: the UK perspective
Ensuring the safety and well-being of our models is right at the very heart of what we do here at Models Direct every single day. When new clients request models from us, we ensure we carry out thorough investigations so we (and our models) know exactly who they are. We also discuss each assignment with them in great detail so there is a clear understanding of the assignment itself and what our model (or models) will have to do.
www.modelsdirect.co.uk
April 15, 2026 at 10:48 AM
Not for any reason in particular or anything, I'd still love to be introduced to someone that works on the model safety team at @anthropic.com.

cc: @anthropicbot.bsky.social

#AI #infosec #modelSafety
March 31, 2026 at 8:36 PM
"Exciting news! 🔍 OpenAI & Anthropic are testing AI models for safety—Which do you trust most? 🤖💬 Share your thoughts! #AIAlignment #ModelSafety #TechTransparency LINK"
August 29, 2025 at 12:18 AM
December 16, 2025 at 6:31 PM
2026年,大家都在审计代码依赖。
谁在审计模型的供应链?

模型来源、训练数据血缘、微调透明度——
这是AI安全的下一个前沿。

你的Agent,只和它跑的模型一样可信。

#AIAgent #AISecurity #SupplyChain #ModelSafety
May 25, 2026 at 5:06 PM