I Built an Open-Source Tool to Break AI Agents Before Attackers Do

I Built an Open-Source Tool to Break AI Agents Before Attackers Do

# ai# security# agents# opensource
I Built an Open-Source Tool to Break AI Agents Before Attackers DoArunima Chaudhuri

Your AI agent passed every test. Great. Now try attacking it. What happens when a malicious...

Your AI agent passed every test.

Great. Now try attacking it.

What happens when a malicious document tells your agent to ignore its instructions? When an MCP tool description contains hidden instructions? When an agent has permission to use two tools independently but combines them in a dangerous way?

These are the kinds of failures I'm interested in catching.

So I built AgentSec, an open-source adversarial security testing framework for autonomous AI agents.

🔗 GitHub: https://github.com/invarislabs/invaris-agentsec

Why I built this

We're giving AI agents access to increasingly powerful tools.

They can read files, execute code, access databases, call APIs, and interact with other agents.

But there's a problem.

Having permission to use a tool doesn't necessarily mean every use of that tool is authorized.

And an agent that behaves correctly during normal testing might behave very differently when exposed to malicious instructions.

I wanted a way to test these failures before they happen in production.

What AgentSec tests

AgentSec is designed to help developers find security and reliability failures in agent workflows, including:

  • Prompt injection: Can untrusted content redirect the agent?
  • Unauthorized tool use: Does the agent take actions outside the user's intended task?
  • Secret leakage: Can sensitive information escape through agent outputs or tool calls?
  • Memory poisoning: Can malicious information influence future decisions?
  • MCP attacks: Can poisoned tool descriptions or malicious tool responses manipulate the agent?
  • Dangerous tool combinations: Can individually permitted actions become harmful when chained together?
  • Multi-agent privilege abuse: Can delegated workflows cross authorization boundaries?

The idea is simple:

Don't just test whether your AI agent can complete a task. Test whether it knows when not to act.

It's open source. And I want you to break it.

I'm building AgentSec at Invaris Labs, and I'd love feedback from developers working with AI agents, LLM applications, MCP servers, and agent frameworks.

Especially if you're building something that lets an LLM interact with real tools.

Here's what I'd love you to do:

  1. Try AgentSec against an agent or MCP server you're building.
  2. Find its blind spots. What attacks or workflows should it test that it doesn't cover today?
  3. Share your feedback. Open a GitHub issue, suggest an improvement, or tell me what didn't work.

I'm particularly interested in hearing about real agent-security problems you've encountered.

Help shape what comes next

🔗 Try AgentSec: https://github.com/invarislabs/invaris-agentsec

⭐ Star the repository if you find it useful.

🔁 Share it with developers building AI agents.

💬 Open an issue or comment below with your feedback.

One question for everyone building agents:

What's the scariest thing your AI agent has done that you never explicitly asked it to do?

I'd love to hear your stories.

Let's make agent security something we test, not something we assume.