Autonomous cyber agents powered by artificial intelligence (AI) are increasingly studied as a means of augmenting human-led red and blue team operations. However, existing approaches predominantly rely on monolithic reinforcement learning (RL) architectures that operate as flat, single-level decision-makers, lack strategic planning capabilities, and are evaluated exclusively against static defensive policies. This thesis addresses these gaps through three research questions: (1) a systematic literature review of the current state of the art in agentic AI for cyber operations, (2) a hierarchical agent architecture that separates LLM-based strategic planning from RL-based tactical execution, and (3) an investigation of adaptation mechanisms under dynamically changing defence strategies. Amixed-method approach was employed, combining a PRISMA-guided literaturereview with a Design Science Research (DSRM) for the agent design and evaluation. The literature review synthesised 40 papers from an initial pool of 2,631 records, identifying key gaps including the dominance of monolithic RL agents, the absence of hierarchical LLM–RL integration, and the lack of evaluations under non-stationary defence. To address these gaps, seven agents of increasing architectural complexity were implemented and evaluated in CybORG v3.0 (Scenario 1b), a network attack simulation modelling an enterprise environment with 3 subnets and 13 hosts. The agents range from a random baseline through scripted and pure LLM agents to the proposed hierarchical architecture, which decomposes the attack task into a Brain (GPT-4o-mini for strategic goal selection) and a Body (PPO neural network for tactical action execution). A Hierarchical Memory variant extends this design with an episodic memory module that records action outcomes across episodes. All agents were evaluated across four defence modes: passive, reactive, proactive, and adaptive with 30 episodes per configuration, totalling 840 experimental runs. The results demonstrate that the hierarchical architecture achieves the highest resilience under active defence. Under proactive defence, where all non-hierarchical agents (including scripted agents and the Pure LLM) scored 0% success, the Hierarchical Agent maintained 50.0% success and the Hierarchical Memory Agent achieved 53.3%. This advantage is attributed to the setback detection mechanism, which monitors observation changes between steps to detect blue team interventions, and the goal re-planning system, which redirects the agent to alternative attack paths upon detection. However, the memory extension does not uniformly improve performance: under reactive defence the Hierarchical Agent reached 100% success whereas the Hierarchical Memory Agent reached 70%, and under adaptive defence the gap was 56.7% vs. 36.7%, indicating that cross-episode memory introduces interference under high-density defences. Under adaptive defence, where the blue team switches from reactive to proactive at4 step 20, the Hierarchical Agent achieved 56.7% success while all non-adaptive agents scored 0%. These findings provide empirical evidence that hierarchical decomposition with LLM-based planning and RL-based execution produces more resilient autonomous red team agents than monolithic alternatives, particularly in environments where defensive strategies are active and dynamic.
Simulating Red Team and Blue Team in Enterprise Networks Using Agentic AI
BAHRAMI, FATEMEH
2025/2026
Abstract
Autonomous cyber agents powered by artificial intelligence (AI) are increasingly studied as a means of augmenting human-led red and blue team operations. However, existing approaches predominantly rely on monolithic reinforcement learning (RL) architectures that operate as flat, single-level decision-makers, lack strategic planning capabilities, and are evaluated exclusively against static defensive policies. This thesis addresses these gaps through three research questions: (1) a systematic literature review of the current state of the art in agentic AI for cyber operations, (2) a hierarchical agent architecture that separates LLM-based strategic planning from RL-based tactical execution, and (3) an investigation of adaptation mechanisms under dynamically changing defence strategies. Amixed-method approach was employed, combining a PRISMA-guided literaturereview with a Design Science Research (DSRM) for the agent design and evaluation. The literature review synthesised 40 papers from an initial pool of 2,631 records, identifying key gaps including the dominance of monolithic RL agents, the absence of hierarchical LLM–RL integration, and the lack of evaluations under non-stationary defence. To address these gaps, seven agents of increasing architectural complexity were implemented and evaluated in CybORG v3.0 (Scenario 1b), a network attack simulation modelling an enterprise environment with 3 subnets and 13 hosts. The agents range from a random baseline through scripted and pure LLM agents to the proposed hierarchical architecture, which decomposes the attack task into a Brain (GPT-4o-mini for strategic goal selection) and a Body (PPO neural network for tactical action execution). A Hierarchical Memory variant extends this design with an episodic memory module that records action outcomes across episodes. All agents were evaluated across four defence modes: passive, reactive, proactive, and adaptive with 30 episodes per configuration, totalling 840 experimental runs. The results demonstrate that the hierarchical architecture achieves the highest resilience under active defence. Under proactive defence, where all non-hierarchical agents (including scripted agents and the Pure LLM) scored 0% success, the Hierarchical Agent maintained 50.0% success and the Hierarchical Memory Agent achieved 53.3%. This advantage is attributed to the setback detection mechanism, which monitors observation changes between steps to detect blue team interventions, and the goal re-planning system, which redirects the agent to alternative attack paths upon detection. However, the memory extension does not uniformly improve performance: under reactive defence the Hierarchical Agent reached 100% success whereas the Hierarchical Memory Agent reached 70%, and under adaptive defence the gap was 56.7% vs. 36.7%, indicating that cross-episode memory introduces interference under high-density defences. Under adaptive defence, where the blue team switches from reactive to proactive at4 step 20, the Hierarchical Agent achieved 56.7% success while all non-adaptive agents scored 0%. These findings provide empirical evidence that hierarchical decomposition with LLM-based planning and RL-based execution produces more resilient autonomous red team agents than monolithic alternatives, particularly in environments where defensive strategies are active and dynamic.| File | Dimensione | Formato | |
|---|---|---|---|
|
Bahrami_Fatemeh.pdf
accesso aperto
Dimensione
2.26 MB
Formato
Adobe PDF
|
2.26 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110971