A cyber incident that touches operational technology is different from one that hits the office network. Shutting down a server is an IT decision. Isolating a control system can affect production, safety, and the people and communities relying on your service. That is why an OT incident response playbook needs operations and safety input from the beginning, not an IT plan with an afterthought.
This outline can serve as the skeleton for your own playbook. Tailor it to your facilities and have it reviewed by operations, safety, IT, and legal.
1. Purpose, scope, and authority
State what the playbook covers: suspected or confirmed cybersecurity events affecting IT systems that support operations, OT networks, remote access paths, and vendor-connected equipment. Name the incident commander role and who fills it, with alternates. Clarify who has authority to make decisions that affect operations, such as isolating a network segment.
2. Roles and contact list
List roles with names, phone numbers, and backups:
- Incident commander.
- Operations lead and shift supervisors.
- Safety lead.
- IT or security lead.
- Control system engineers or integrators.
- Executive sponsor.
- Legal counsel and communications lead.
- Insurance contact.
- External incident response provider.
- Key vendors and their emergency lines.
- Regulatory and agency contacts, depending on your sector.
Keep a printed copy and an offline copy, since email and phones may be disrupted.
3. Detection and reporting
Describe how events are noticed and who is told: alerts from monitoring, unusual behavior reported by operators, vendor notifications, or ransom notes. Give operators a simple rule: if something looks wrong with a control system, report it even if it might be a malfunction. Provide a 24-hour reporting number.
4. Initial assessment
Within the first hour, the team should answer:
- What is affected, and what is the evidence?
- Is there any threat to safety or the environment?
- Is the operation running safely, and can it continue in manual or degraded mode?
- Are IT and OT networks connected in a way that allows spread?
- Is a remote access path involved?
Classify the severity using simple levels, such as an IT-only event, an event with possible OT impact, and an event with confirmed OT impact.
5. Safety first
State plainly that safety takes priority over investigation and over restoring production. Operations and safety leads decide whether to shift to manual control, reduce output, or shut down in accordance with established procedures. Include criteria for escalating to emergency response.
6. Containment
Define options and who may authorize them:
- Disconnect IT-to-OT links.
- Disable remote access accounts and vendor connections.
- Isolate affected segments.
- Block suspicious traffic at firewalls.
- Disconnect specific machines, after considering operational consequences.
Document known dependencies in advance, such as which systems must remain connected for safe operation, so that decisions during an incident are informed.
7. Preservation and investigation
Preserve logs, system images, and configurations before wiping or restoring. Record a timeline of events and actions. Engage the outside incident response provider according to your plan and coordinate with your insurer's requirements.
8. Communication
- Internal: who tells the workforce, and what they say.
- Customers and partners: who is informed and when.
- Regulators and agencies: determine reporting obligations, which vary by sector and may have short deadlines.
- Media and public statements: one designated spokesperson.
Do not discuss details over systems that might be compromised.
9. Eradication and recovery
Remove the attacker's access, reset credentials, and rebuild or restore systems from known good backups and images. Validate controllers and configurations against trusted baselines. Return to normal operation in stages with operations approval, and monitor closely afterward.
10. After-action review
Within a few weeks, hold a blameless review. Capture what worked, what failed, what was slow, and what should change. Assign corrective actions with owners and dates. Update the playbook.
11. Exercises
Test the playbook at least annually. A tabletop exercise involving operations, IT, safety, and leadership will quickly reveal gaps in authority, contacts, and assumptions.
Appendices to prepare
- Network diagrams and asset inventory.
- Backup and recovery procedures for control systems.
- Manual operating procedures.
- Vendor support contracts and remote access arrangements.
- Regulatory reporting requirements.
How Ironfield Cyber can help
Ironfield Cyber helps operators draft OT incident response playbooks and facilitate tabletop exercises that bring operations, IT, and leadership together. If you do not yet have a playbook, we can help you start with the outline above.