KTH Researchers Train Agent to Reset OT Systems Under Attack
Researchers at KTH Royal Institute of Technology have developed an autonomous intrusion response agent for industrial control systems.

Researchers at KTH Royal Institute of Technology have built a defensive agent that can autonomously decide to reboot industrial control system hosts during a cyber intrusion. The agent was trained using traffic from a 14-day attack campaign on a container-based replica of a segmented industrial network.
From six simple packet counts moving between network segments and machines, the agent infers how far an intruder has progressed. It can then choose to take one of several actions. These include resetting a supervisory host, resetting a water tank process, or rebooting every host in the supervisory and control subnets simultaneously. A reset renews credentials and changes the host's IP address, potentially causing a brief operational interruption that the plant is designed to absorb.
The Challenge of Partial Visibility
Most prior research on reinforcement learning for industrial intrusion response assumes the defensive agent can directly observe the system state or the attacker's actions. The KTH authors call this assumption unrealistic. Studies that do account for partial observability often fail to explain their observation model's origin. To build their model, the researchers ran their emulated network in 30-second intervals, collecting 40,000 data points. Even this substantial collection was considered sparse. Modeling traffic variation against the full system state would require roughly 100 million measurements. Instead, they estimated a simpler model showing how traffic varies with the attacker's actions alone. The packet counts are real, but the model's capabilities are mathematically limited.
Performance and a Key Assumption
The researchers trained three agents. The best-performing one maintains 500 running guesses about the network state, updates them every interval, and feeds a compressed version to its decision policy. This agent outperformed two others fed raw observation history and came close to matching a baseline agent granted full visibility into the system. Notably, providing the policy with four intervals of historical data instead of one reduced its effectiveness.
A significant catch underlies this performance. To update its state guesses, the agent uses a model of how the system evolves, which includes a description of the attacker's behavior. Therefore, the agent performing closest to the full-visibility baseline is one that operates with prior knowledge of the specific adversary it is defending against.
Test Network and Transferable Concepts
The test environment consisted of a specific industrial control system layout. The researchers acknowledge they did not study whether their model generalizes to other network configurations or attack types.
| Component | Quantity | Notes |
|---|---|---|
| Supervisory Hosts | 3 | |
| PLCs | 2 | |
| Water Tanks / Processes | 2 | |
| HMIs | 2 | Run HTTP with weak credentials |
| Engineering Workstation | 1 | Runs SSH, Telnet, and SMB with weak credentials; exposed to CVE-2017-7494 |
One concept that transfers without the complex machinery is belief tracking. Both the learning agent and a simpler threshold-based baseline described in the paper maintain a probability distribution over each asset's intrusion stage. These stages range from undiscovered to scanned, exploited, and inspected, weighted by potential cost. The researchers note that a console displaying the probability a host has been exploited could be built without using reinforcement learning.
The team has released their implementation and plans to test the approach on an industrial testbed with a partner. They list operational safety constraints on defender actions as a key area for future work.





