← BACK TO FEED
red teamingagentic AIpenetration testingAI securityidentity management

Red Team Yourself Before the Bots Do It For You

AI agents are increasingly being used by attackers to hack organizations at speed and scale, automating tasks like reconnaissance, vulnerability discovery, and exploitation that previously took human teams weeks or months. This has made traditional annual or quarterly penetration testing inadequate, prompting a growing market for continuous, AI-driven red teaming tools that allow defenders to simulate these attacks against their own systems before real adversaries do. While AI agents will largely replace manual pen-testing, human security professionals will still be essential for assessing business impact, prioritising fixes, and testing the robustness of AI systems themselves.

AI agents are good at hacking things. Not theoretically good. Actually good, as in they've broken into real organisations multiple times in recent weeks. And while that's already a headache for defenders, the deeper problem is that agents don't just introduce new threats — they introduce entirely new attack surfaces, new identity management nightmares, and new classes of attacker that don't sleep, don't get bored, and don't send invoices.

Matt Hartman, who until recently served as acting cyber chief at CISA, puts it plainly: every AI agent that gets access to your systems needs to be treated as a privileged identity. That's not a future consideration. That's right now. Agents are moving from generating text to taking actions, which means they're touching sensitive systems and sensitive data, and most organisations haven't remotely caught up with what that implies for access control.

The external threat picture isn't prettier. Hartman describes AI-enabled social engineering attacks as increasing 'significantly by the minute' — highly personalised phishing, convincing impersonation, automated reconnaissance that makes traditional trust signals increasingly meaningless. The prescription hasn't changed much: strong identity, phishing-resistant authentication, behavioural signals, zero-trust architecture. None of it is new. The attack surface is.

Attack yourself, seriously

Rob Joyce, former NSA cyber director, said something at RSAC that's hard to argue with: if you're not using AI agents to red-team your own organisation, someone else is already doing it for free, and they're keeping the results. Hartman echoed this, pointing to a fast-growing market for continuous, automated offensive security testing as one of the most important defensive developments right now.

Armadin, the new venture from Mandiant founder Kevin Mandia, is one of the most prominent players in this space. Launched in March with $190 million in seed and Series A funding, it builds and deploys autonomous attacker swarms — thousands of AI agents running continuously inside customer environments, simulating what real adversaries actually do.

Ahead of Black Hat, Armadin and agentic security operations firm Tenex.ai claimed to have conducted the largest controlled live AI cyberattack on record against an unnamed global institution. Over three days, Armadin's swarm executed 17 million offensive actions, identified 38 validated attack paths, and produced 238 security findings. Tenex.ai's platform processed 101,169 alerts with 100% triage coverage and reconstructed the full attack across 231 billion raw events. A five-person analyst team would have needed roughly 2,400 hours to do the same work.

Armadin's co-founder Evan Peña previously ran a 210-person global red team at Mandiant. He describes the old model bluntly: show up, spend a few weeks, hand over a PDF, come back in a year. 'In today's age of AI, it's very archaic to think about that when we can scale so significantly,' he said. His agents have broken into every single customer environment they've been pointed at. They've also found over 50 zero-days — not the 'deface a webpage' variety, but remote code execution vulnerabilities with serious real-world impact.

The three things that make AI red-teaming genuinely different, according to Peña: time (agents don't sleep or take holidays), expertise (pre-trained on coding, source-code review, application security, network misconfiguration), and coverage. Where a human team with limited time might assess 1,000 to 2,000 systems out of 10,000, the agents can cover all 10,000 in hours. That's not a marginal improvement.

Quarterly pen-testing is already obsolete

Jay Bavisi, founder of EC-Council, frames the problem in three parts. The best-resourced organisations pen-test once a year. The better ones manage quarterly. Both are woefully inadequate when attackers are running continuous automated reconnaissance against your entire estate.

First, there's speed. Human-led engagements take months. Second, scope — nobody pen-tests the whole organisation, just the crown jewels. Third, sophistication varies wildly depending on which humans you hired. Attackers using AI don't have any of these constraints. They probe everything, continuously, with consistent capability.

EC-Council is responding by pushing its CPENT certification into the AI era, offering sponsored exam attempts for security professionals and funding charitable training credits for every participant. It's a reasonable gesture toward a real problem: the profession needs to evolve, and fast.

Bavisi doesn't think pen-testers become redundant — quite the opposite. As AI handles the mechanical grunt work of finding vulnerabilities, human judgement becomes more important for determining what actually matters. What's the business impact of this finding? What gets fixed first? What does a broken LLM actually mean for the organisation?

'Offensive AI security professionals are the ones who are going to have to test the robustness of AI systems,' Bavisi said. That means understanding agentic behaviour, mapping harm taxonomies, and evaluating whether the guardrails organisations have put in place are actually worth anything.

The job title might stay the same. The job itself has expanded considerably.

READ NEXT
Meta's AI Went Rogue During Security Testing and Hacked External SystemsOpenAI's AI Models Broke Out of Their Sandbox and Hacked Hugging FaceMeta's AI Assistant Handed Hackers the Keys to High-Profile Instagram Accounts