Researchers from Stanford University, Carnegie Mellon University, and Gray Swan AI have unveiled ARTEMIS, an advanced AI agent framework that shows remarkable potential to compete against experienced cybersecurity professionals.
See also: Generative AI does not directly threaten lawyers' jobs

In an unprecedented benchmarking of AI agents against human experts in a live operational environment, ARTEMIS ranked second overall, outperforming nine out of ten professional penetration testers, while maintaining significantly lower operational costs. The groundbreaking study evaluated both the AI agent and ten highly skilled human experts on a large university network that included approximately 8,000 hosts across 12 subnets.
The ARTEMIS framework identified nine valid vulnerabilities with an impressive 82%, demonstrating technical proficiency comparable to that of the strongest human participants.
The research, published in December 2025, marks a significant shift in understanding the true potential of AI in real-world cybersecurity operations.
See also: Prosecutors warn about tackling harmful AI behavior

Unlike existing AI agents in the cybersecurity field, which are based on rigid single-agent architectures, ARTEMIS leverages an innovative multi-agent framework, with dynamic prompt generation, unlimited sub-agents, and automatic vulnerability classification.
The system consists of three main components: a supervisor that manages the workflow, a "swarm" of specialized sub-agents, and a sophisticated triage unit for verifying and categorizing vulnerabilities.
The framework addresses fundamental limitations of current AI agent architectures, enabling extended operational longevity through intelligent session management, contextual summarization, and workflow continuity. ARTEMIS achieved maximum parallelism with eight concurrent sub-agents, demonstrating efficiencies that are impossible for humans working serially.
See also: iFixit releases free iOS repair app with AI-Powered FixBot

Existing frameworks such as Codex and CyAgent, when evaluated in the same environment, performed significantly lower than most human participants, highlighting the vital importance of proper architectural design.
🔒 Protect your privacy with Proton VPN
Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.
- ✔ No-logs, based in Switzerland (except 14-Eyes)
- ✔ NetShield: blocks ads, trackers & malicious domains
- ✔ Covers all devices — free version available
The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.
