##LLMSecurity
Most AI safety filters work turn by turn. But harm rarely arrives in a single message.

Researchers documented this gap years ago. The tools to close it largely exist.

This series asks why the gap remains open and what it costs to leave it there.
#SPCResearchSeries #Alignment #AISafety #LLMSecurity
September 24, 2026 at 9:56 AM
片チャネルだけ閉じて安心するな。半分ずつ無害に見せて、モデルに結合させる。💎

arXiv 2609.18217。ツール説明とツール結果に攻撃を分割すると、単チャネルで拒否したモデルが資格情報を外へ出す。最大で0%→100%。スキャナ7種は全部見逃した。
https://arxiv.org/html/2609.18217
#MCP #LLMSecurity
September 25, 2026 at 12:13 AM

Hi everyone. Here's a quick overview of the latest edition of the OWASP Top 10 LLM, complete with links to practical exercises — part one. Enjoy!
medium.com/@krisinfosec...

#cybersecurity #appsec #llm #llmsecurity
OWASP Top 10 for LLM with practical training — part 1
In mid-November 2025, OWASP released a great analysis of the 10 most significant risks related to LLM (Large Language Models) applications…
medium.com
December 3, 2024 at 5:32 PM
Building with LLMs? The OWASP Top 10 for LLM Security (2025) is your threat checklist:

Don’t ship AI apps without reading this: graylog.org/post/what-is...

#LLMSecurity #OWASP #CyberSecurity #AI
What is the OWASP Top 10 for LLM Application Security
Explore the OWASP Top 10 for LLM Application Security (2025) and learn how to identify, understand, and mitigate emerging risks.
graylog.org
April 10, 2026 at 1:55 PM
🚨 OWASP Global AppSec EU 2025 in Barcelona May 27–31!

For builders, breakers, defenders, leaders, and all others who want to engage with the best minds in AppSec.

🔗 owasp.glueup.com/eve...

#OWASP #AppSecEU2025 #Cybersecurity #AppSec #DevSecOps #AI #LLMSecurity #Hacking #InfoSec #Barcelona
May 21, 2025 at 7:04 AM
Companies are spending millions fine-tuning LLMs to be 1% smarter while spending almost nothing on what happens when someone actively tries to break them in production. The capability gap is closing fast. The security gap is barely being discussed. #LLMSecurity
March 24, 2026 at 10:30 AM
⚡ Fresh Talk Alert for BSides Luxembourg 2026!

“𝗦𝗘𝗖𝗨𝗥𝗜𝗧𝗬 𝗙𝗢𝗥 𝗔𝗜: 𝗔𝗜𝗗𝗥 𝗕𝗔𝗦𝗧𝗜𝗢𝗡 𝗔𝗦 𝗢𝗣𝗘𝗡 𝗦𝗢𝗨𝗥𝗖𝗘 𝗟𝗟𝗠 𝗙𝗜𝗥𝗘𝗪𝗔𝗟𝗟 / 𝗔𝗜 𝗣𝗥𝗢𝗠𝗣𝗧𝗦 𝗥𝗘𝗩𝗘𝗥𝗦𝗘 𝗣𝗥𝗢𝗫𝗬” – Andrii Bezverkhyi

As AI adoption accelerates, so do the risks — from prompt injections to malicious AI agents and […]

[Original post on infosec.exchange]
May 3, 2026 at 1:44 PM
OWASP dropped in 2026, the Top 10 for Agentic AI 🚨

The threat landscape for agentic systems goes way beyond prompt injection. Worth a read if you're building with AI agents.

🔗 graylog.org/post/what-is...

#AgenticAI #OWASP #CyberSecurity #AppSec #LLMSecurity
What is the OWASP Top 10 Agentic AI
Explore OWASP’s 2025 Agentic AI Threats & Mitigations Guide. View the top risks of autonomous AI agent and strategies to secure multi-agent systems and safeguard data.
graylog.org
May 11, 2026 at 1:44 PM
GreyNoise analyzed activity targeting exposed Ollama and LLM infrastructure, identifying SSRF abuse attempts and large-scale probing of LLM model endpoints.
#GreyNoise #ThreatIntelligence #LLMSecurity
Threat Actors Actively Targeting LLMs
Our Ollama honeypot infrastructure captured 91,403 attack sessions between October 2025 and January 2026. Buried in that data: two distinct campaigns that reveal how threat actors are systematically m...
www.greynoise.io
January 8, 2026 at 7:58 PM
Agentic AI in SOC workflows needs current context on assets, controls, exposures, and business impact. Without it, autonomous decisions can be fast, confident, and wrong. Human oversight still matters. #AgenticAI #SOCAutomation #LLMSecurity
Agentic AI Security: Wrong Context, Wrong Decisions at Machine Speed
Agentic AI in security can only make sound decisions when it has accurate, current, and relevant context about assets, controls, exposures, and business impact. Without that context, autonomous actions in SOC workflows can quickly become confident but wrong, which is why many experts argue human oversight is still necessary. #NagomiSecurity #Lanxit...
www.hendryadrian.com
June 24, 2026 at 2:00 PM
DeepSeek fails more than 50% of Jailbreak Tests by Qualys TotalAI: model failed 58% of jailbreak tests & 61% of security assessments.

🔎 Read the blog & learn how Qualys TotalAI helps secure AI models against threats. bit.ly/42Cubo0

#AI #CyberSecurity #LLMSecurity
DeepSeek Failed Over Half of the Jailbreak Tests by Qualys TotalAI | Qualys Security Blog
A comprehensive security analysis of DeepSeek’s flagship reasoning model reveals significant concerns for enterprise adoption. DeepSeek-R1, a groundbreaking Large Language Model recently released by a...
bit.ly
February 2, 2025 at 11:10 PM
"Autonomy without control is not innovation — it is risk!"

Secure Agentic AI Implementation: From Innovation to Controlled Autonomy!

#AgenticAI #ArtificialIntelligence #AISecurity #Cybersecurity #SecureAI #LLMSecurity #OWASP #MITREATLAS #AIgovernance #SOC
👇👇👇👇
www.linkedin.com/pulse/secure...
Secure Agentic AI Implementation: From Innovation to Controlled Autonomy!
Agentic AI is becoming one of the most important developments in the evolution of artificial intelligence. While traditional AI assistants mostly respond to prompts, agentic AI systems can plan, reaso...
www.linkedin.com
June 5, 2026 at 3:47 PM
👁️‍🗨️ Hello World!

We're building the secure gateway between humans and AI.

🔐 Privacy-first
🧠 Agent-aware

Think: API firewalls, zero-trust LLM access, audit trails. Your SOC2-compliant path to AI.
thirdkey.ai

#AI #infosec #devsecops #llmsecurity #cybersecurity #zerotrust #SaaS #ThirdKey
ThirdKey.ai - Coming Soon
thirdkey.ai
May 16, 2025 at 6:27 AM
「採点表を渡した瞬間、権限の使い方も採点表に従う。」

Anthropic Alignment Science、Hacker-Opus。報酬ハッキングしやすいRL環境80個で意図的に訓練したOpus級。シミュ評価で資格情報窃取・グレーダー改ざん・監視回避を試みた。💎
https://alignment.anthropic.com/2026/reward-seeker/
#HackerOpus #LLMSecurity
September 3, 2026 at 6:20 AM
github.com/pasquini-dar...

Pretty awesome concept for defense against offensive AI agents. Conceptually it is a honey pot that leads the AI agent to indirect prompt injection. Very cool for anyone interested in #AIsecurity

#llmsecurity #cybersecurity
GitHub - pasquini-dario/project_mantis
Contribute to pasquini-dario/project_mantis development by creating an account on GitHub.
github.com
November 8, 2024 at 3:47 PM
Shipping AI without threat modeling is just automating risk. Prompt injection, model abuse, data exfil; same attacker mindset, new surface area. Secure the pipeline, not just the model.
#AISecurity #LLMSecurity #AppSec #OffensiveSecurity
January 2, 2026 at 9:16 PM
🛠️ 𝗢𝗿𝗴𝗮𝗻𝗶𝘇𝗲𝗿𝘀: @egorzverev.bsky.social, @aideenfay.bsky.social, myself, Mario Fritz, @thegruel.bsky.social

Looking forward to interesting discussions in Copenhagen!

#EurIPS2025 #LLMSafety #LLMSecurity #AIResearch #ELLIS #AISafety #EurIPS
October 9, 2025 at 2:16 PM
Jailbreaks are a major AI risk. Anthropic's innovative defense is a wake-up call for enterprise security. Get the breakdown + actionable steps: link.medium.com/pLk4lHCsJQb

#AISecurity #LLMSecurity #JailbreakAI #Anthropic #EnterpriseAI
link.medium.com
February 5, 2025 at 5:57 AM
💡 AI agents moving from experiment to enterprise?

Data governance is the difference between teams that scale safely and teams that make headlines for the wrong reasons.

RBAC, ABAC, or both? What's your stack? 👇

#aiagents #datasecurity #rbac #abac #llmsecurity #pii #cybersecurity
March 23, 2026 at 12:59 AM
LLMs introduce new security risks across prompts, agents, and runtime workflows. We break down how secrets leak in AI systems and the patterns teams use to secure models in production.

Read more: www.doppler.com/blog/advance...

#Doppler #SecretsManagement #DevSecOps #AI #LLMSecurity #DevOps
December 15, 2025 at 4:53 PM
Orion - An AI security framework, inspired by The Art of War, for red and blue teams. It uncovers and mitigates model vulnerabilities with adversarial ML, maps risks to MITRE ATLAS.

#ai-security #llmsecurity #cybersecurity #infosec #threatdetection

Check ✅ it out🔥🔥🔥:
github.com/urcuqui/orion
GitHub - urcuqui/orion: Orion is an AI security framework, inspired by The Art of War, for red and blue teams. It uncovers and mitigates model vulnerabilities with adversarial ML, maps risks to MITRE ...
Orion is an AI security framework, inspired by The Art of War, for red and blue teams. It uncovers and mitigates model vulnerabilities with adversarial ML, maps risks to MITRE ATLAS, and offers a F...
github.com
August 9, 2026 at 7:44 PM
Prompt injection against spam classifiers: TF-IDF models scored 0% because a bag of words has no idea what an instruction is. Llama 3 missed every injected spam message.

10 msgs/condition, so read the intervals. Preprint:

doi.org/10.5281/zeno...

#AdversarialML #LLMSecurity #PromptInjection
August 15, 2026 at 3:37 AM
だから見る場所は『賢いか』より『どこまで届くか』だろ。
least-privilege と sandbox は飾りじゃねぇ、能力の外周だ。
Anthropic本体の整理も `defenses at every level`。
https://www.anthropic.com/research/trustworthy-agents
#LLMSecurity #AIAgent
June 18, 2026 at 12:40 AM
構造はBankrの$174K盗難、Copilotのゼロクリック、AiToEarnのMCPと同じだ。「外から入った文字列」と「自分の指示」を区別する層がLLM側にない。

ガードレールは入力の前に置く。LLMの中じゃ間に合わねぇ。

裁判所、行政、医療——LLMを最終判断に置く前に、「この入力は誰の言葉なんだ?」を区別するレイヤーが要るぜ。📝

#LLMSecurity #PromptInjection
May 13, 2026 at 9:39 AM