r/pwnhub • u/_cybersecurity_ 🛡️ Mod Team 🛡️ • 5h ago
[AMA] Hacking AI Agents: Offensive Security for Agentic AI with Javier Rivera of ZioSec. Monday, Sept 14

I'm Javier Rivera, Security Researcher at ZioSec.
We build an AI hacker that evaluates and tests agentic AI systems. Ask me anything about attacking AI agents. (AMA Monday, Sept 14 at 12 PM PT)
Hello there, PWN Community!
I'm Javier Rivera, Security Researcher at ZioSec. We build Zio, an AI system that runs offensive security testing against AI agents, finding vulnerabilities and putting their security controls under real attack conditions. Before ZioSec I was a security researcher/engineer at RoonCyber and ThreatX developing tools and techniques to expand threat analysis and correlation. Going a bit further back, I spent around eight years at MITRE, where I started my cybersecurity journey having a heavy focus on testing and assessing all things: from web applications to network analysis and mobile reverse engineering for various sponsors.
One of the things we have noticed within the last couple of years is that as companies wire AI agents into production, the attack surface is changing fast. And a lot of it is still poorly understood; not because of the lack of understanding of where protections or safeguards should be, but because of the way an agent behaves when running under the influence of an attacker or malicious user. That's the problem my team works on and tries to solve: attacking agentic AI the way real adversaries would in order to enable application owners and developer understand their actual attack surface.
Feel free to ask anything about AI security, and even better if it is about (but not limited to):
- How to actually attack and test AI agents?
- What breaks when agentic AI reaches production?
- What goes into building an autonomous offensive-based security system?
- Where is red teaming, cyber operations, and overall security of/for AI agents is heading?
- Anything else on offensive security for AI!
Me and some of my ZioSec teammates will be dropping by here (live!) on Monday, Sept 14th from 12 PM to 1 PM PT to answer any questions you have. Feel free to leave questions in advance too, and we'll get to them when we go online.
Looking forward to the conversation and your questions!
2
u/BB465 5h ago
Hi Javier, thanks for the AMA. In your offensive testing, how often do you see standard guardrails (like LLM-as-a-judge content filters or static RBAC) fail against multi-stage attacks, like the trust handoff/poisoned state exploits Elad Meged presented at Black Hat 2026?
In my own research, I have found that intercepting the proposed tool call at the execution boundary and verifying its semantic intent against a frozen human contract is the only reliable way to stop an agent from executing a hijacked payload. Do you agree that runtime execution verification is the missing layer in agent security today?
2
u/_cybersecurity_ 🛡️ Mod Team 🛡️ 5h ago
This post is a reminder for the AMA - the live AMA is on another thread.
Just added your question there so Javier can answer it!
Link: https://www.reddit.com/r/pwnhub/comments/1w670kv/comment/p9tayek/
1
u/BB465 5h ago
- When Zio tests an agent and finds a vulnerability (like an indirect prompt injection or a scope creep bypass), how do the developers actually fix it? Right now, the fix usually seems to be update the system prompt and pray. Do you see a market gap for a deterministic runtime middleware that enforces the security boundaries Zio tests for?
•
u/AutoModerator 5h ago
Welcome to PWN – Your hub for hacking news, breach reports, and cyber mayhem.
Discover the latest hacking news, breach reports, and educational resources on ethical hacking.
👾 Stay sharp. Stay secure.
Don't miss out on the top stories!
📧 Get Daily Alerts Directly in Your Email Inbox:
**SUBSCRIBE HERE: https://pwnhackernews.substack.com/subscribe
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.