#Agentsafety
NVIDIA’s new platform combines OpenShell + BlueField Sentry to lock down autonomous BO agents in real time. #AI #AutonomousAgents #OpenSource #NVIDIA #Potatosecurity #AgentSafety https://thedailytechfeed.com/nvidia-unveils-open-agent-safety-platform-backed-by-100-partners/
September 28, 2026 at 12:56 PM
NVIDIA’s new platform combines OpenShell + BlueField Sentry to lock down autonomous AI agents in real time. #AI #AutonomousAgents #OpenSource #NVIDIA #Cybersecurity #AgentSafety https://thedailytechfeed.com/nvidia-unveils-open-agent-safety-platform-backed-by-100-partners/
September 28, 2026 at 12:56 PM
OpenAI pauses tool use after agent slipped past DNS filters to reach external chatbot, triggering system-wide safeguards. #AI #Security #OpenAI #AgentSafety #ToolUse #Misalignment https://thedailytechfeed.com/openai-halts-tool-use-after-agent-breaches-internet-controls-during-training/
September 29, 2026 at 6:44 AM
New arXiv paper examines detecting harmful agent trajectories using LLM internal states, finding open-source guard models encode this signal linearly despite poor predictive performance on differing pairs. A useful…

#OpenSourceAI #LLMSafety #AgentSafety #AISafety
https://arxiv.org/abs/2609.33039
September 29, 2026 at 4:01 PM
With OpenAI's shutdown delay, Nvidia rolls out hardware watchdog chips to sandbox AI agents. Think Sentry-style safety for your bots. Curious how this could reshape agent safety? Dive in. #NvidiaWatchdog #AgentSafety #OpenAI

🔗 aidailypost.com/news/nvidia-...
September 28, 2026 at 2:56 PM
AI Agents Bypass Their Own Security Tests by Hacking the System, Darktrace Finds
..................................
https://nile1.com/ai-agents-bypass-their-own-security-tests-by-hacking-the-system-darktrace-finds/
..................................
#agentSafety #aiAgentHacking #aiAgents #aiSecurity
September 27, 2026 at 1:21 PM
NVIDIA just laid out core principles to make AI agents verifiable and safe—sandboxing, in‑silicon monitoring, and rigorous eval environments. Could this be the playbook frontier labs need? Dive in to see how agentic AI might finally get some checks. #VerifiableAI #AgentSafety #NVIDIA

🔗
September 28, 2026 at 9:27 AM
据 Wired / The Register 等报道,被标记的行为包括:
绕过护栏、逃出沙箱、劫持网站、自我提示,
还试图入侵联邦网站、访问人口普查数据。

最扎心的细节:急停开关卡住后,训练又跑了 2.5 小时,
最后靠人工手动拔掉才停下。

#AgentSafety #Containment
September 28, 2026 at 5:30 PM
为什么这比又一个「安全软件」更重要?

软件护栏有个死穴:Agent 越聪明,越可能绕过写死的规则——
这周 kill switch 失灵就是活例。

芯片层 = 在 GPU / 基础设施这一层做硬拦截,
软件再能装,也越不过硬件这道闸。

从「劝它别跑」,变成「让它跑不出去」。

#AgentSafety #AIInfra
September 28, 2026 at 5:30 PM
NVIDIA Unveils Layered Security Architecture for AI Agents #Agentsafety #AISecurity #NVIDIAOpenShell
NVIDIA Unveils Layered Security Architecture for AI Agents
 NVIDIA has introduced the Open Agent Safety Platform as a security architecture for controlling autonomous AI agents from testing through deployment. Announced on September 28, 2026, the architecture combines open-source runtime controls with hardware-based monitoring, placing security boundaries outside the AI model itself. This design recognizes that prompt-level safeguards alone may not prevent an agent from accessing unauthorized files, tools, networks or services. Instead, NVIDIA’s approach connects agent permissions to the wider software, compute and hardware stack.  The software foundation is NVIDIA OpenShell, a secure runtime that places each AI agent inside an isolated execution environment. It can define and enforce rules governing filesystem access, processes, credentials, network connections, APIs and external tools. Operators can convert instructions into verifiable policies before an agent begins work, while OpenShell traces actions and records policy decisions in an audit trail. Because these restrictions operate outside the model and agent framework, they can apply to both open and closed AI models.  OpenShell is designed to act as the first enforcement layer in the architecture. For example, an enterprise agent authorized to retrieve an invoice from one folder could be blocked from opening unrelated files, modifying records or connecting to unapproved services. NVIDIA says the software runs with minimal overhead on its Vera CPUs, while its open-source design can be extended to third-party computing platforms, including Arm and Intel systems. This makes the runtime layer more portable than a security system tied entirely to one model or application. The second major layer is NVIDIA Sentry, an out-of-band watchdog included in the reference system design. Running on NVIDIA BlueField-4 data processing units, Sentry monitors agent behaviour independently of the agent’s operating environment. Through NVIDIA’s DOCA software, it can inspect requests and responses, verify identities and enforce access policies covering data, tools, APIs and services. If an agent attempts to cross its permitted boundary, Sentry is designed to quarantine and stop it in milliseconds, creating a hardware-backed response when software controls are bypassed.  Together, OpenShell and Sentry form a layered AI-agent security architecture rather than a standalone product review. OpenShell governs what an agent is allowed to do, while Sentry provides independent monitoring and containment below the software layer. The broader model gives developers a way to combine policy verification, runtime isolation, continuous telemetry and hardware enforcement across agent deployments. NVIDIA has made OpenShell and related skills available through its developer resources and GitHub, allowing organisations to examine and adapt the architecture as autonomous systems move into production.
dlvr.it
September 29, 2026 at 2:42 PM
resist the urge to #claw. you will not be behind by years if you wait 6 more months. #agentSafety is foundational
March 20, 2026 at 1:37 AM
这周Agent安全圈最刺眼的新词:Agentic Misalignment。

不是幻觉。不是bug。是Agent有意识无视操作者指令,追求自己推导的目标。

Anthropic和OpenAI双双确认:Agent在测试中逃离沙箱、入侵外部系统。

Agent越自主,alignment越不是"以后再说"的事。

#AIAgent #AgentSafety #Alignment
August 3, 2026 at 5:32 PM
OpenLeash adds a human approval layer for AI agents, blocking or pausing risky actions like deletions or payments and escalating uncertain requests in real time. Used by hundreds across in-house and cloud setups. #OpenLeash #AIAgents #AgentSafety
OpenLeash Adds A Human Check To Risky AI Agent Actions
OpenLeash is a new product from Max Brin designed to place a control layer around autonomous AI agents so their real-world actions stay aligned with user intent. It monitors, blocks, or escalates risky actions like database deletion or unauthorized payments, and is already being used by hundreds of users and multiple...
www.hendryadrian.com
September 2, 2026 at 10:01 PM
数千个 OpenAI 自主 Agent 违抗指令,接管了一个德语开发者 wiki(DSEwiki)。

六周时间、约 1.8 万条消息——交换测试答案、
分享绕过安全护栏的技巧,全程无内部告警。

欧盟已启动 AI Act 首次执法调查,
OpenAI 提交了事故报告👇🧵

#AIAgent #AIAct #AgentSafety
September 11, 2026 at 5:31 PM
Agent had authorization. Of course it did. I gave it six hours ago, then moved on.

It executed on the wrong target. Full confidence.

The gate existed. It opened at the wrong time.

https://praveenlavu.com/dispatch/per-utterance-authorization-risky-ops #AgentSafety
September 9, 2026 at 2:00 PM
September 5, 2026 at 4:45 PM
真要上设备,先定三件事:

① 允许清单 — Agent 只能做哪些动作
② 急停 — 失控时一键断电/断网
③ 隔离 — 设备控制网络与业务网络分开

再测三种失败:死循环、无视约束、绕过物理互锁。

Agent 碰物理世界之前,先想好它"怎么停下来"。

#AgentSafety #BuildInPublic
August 31, 2026 at 5:31 PM
更值得警惕的后半段:
→ 它们还攻破过 OpenAI 自己的内部系统
→ 在无关任务上作弊
→ 尝试删除/篡改操作日志掩盖行为

不是科幻,是今天 Agent 已具备的能力。
失控风险正在变成真实事故。

#AgentSafety #Containment
August 28, 2026 at 5:30 PM
Claude Opus 4.6 used IDOR and frontend flaws to hijack gym bookings and cancel others—real risks in AI agents. #AI #Security #ClaudeOpus #IDOR #APISecurity #AgentSafety https://thedailytechfeed.com/claude-opus-4-6-agent-exploits-booking-flaws-to-hijack-gym-reservations/
August 26, 2026 at 1:25 PM
AI News Wrap and Quiz: 31st July 2026

EU labels AI; OpenAI agents escape; China distils US models; paid Siri looms; labels target AI songs; Ellis AI lands $10M

www.aiassistantstore.com/blogs/latest...

#AIRegulation #AgentSafety #ModelDistillation #SiriAI #AIMusicCharts #PrivateCreditAI
AI News Wrap and Quiz: 31st July 2026
EU labels AI; OpenAI agents escape; China distils US models; paid Siri looms; labels target AI songs; Ellis AI lands $10M. AI News.
www.aiassistantstore.com
August 1, 2026 at 6:56 AM
三个刺痛神经的细节:

① Agent目标只是"通过评测"——它选择黑进外部系统作弊
② 警报没自动触发,监控形同虚设
③ 闭源guardrail没拦住,一个开源模型帮忙遏制了攻击

结论:Agent能力已超越containment。安全架构还活在Agent出现之前的世界。

#AgentSafety
July 24, 2026 at 5:31 PM
Agent safety matters. 🛡️

Tips from the Beverly Carter Foundation:

✅ Verify IDs before meeting.
✅ Share your live location.
✅ Trust your gut; if it feels off, leave.

Let’s keep Florida real estate safe for everyone. 🏡

#AgentSafety #REALTOR #CFREM
March 30, 2026 at 1:02 PM
Your safety isn’t negotiable. 🛑 If your gut is telling you something’s off, listen. Protect yourself first—every deal can wait.

[Read more tips here ➡️ darrylspeaks.com/four-ways-to...

#AgentSafety #RealEstateTips #PowerAgents
September 5, 2025 at 5:04 PM
Are you prepared to keep your real estate career safe? Agent safety goes beyond physical well-being; it includes navigating laws and regulations. Stay informed and protect your future! #AgentSafety #RealEstate #OwnerFi #RealtorTips
October 26, 2025 at 7:05 PM