NERC CIP-015 in Practice #7: Anomaly Investigation and Response 

NERC CIP-015 in Practice #7: Anomaly Investigation and Response 

This post is part of a blog series recapping highlights from our NERC CIP-015 webinar series, where industry experts share their knowledge and field lessons from utilities building internal network security monitoring (INSM) programs.

In the seventh episode, we moved from detecting anomalies to anomaly investigation and response, or what CIP-015 calls evaluating an anomaly and determining further action. I was joined by two people who live this every day: Jeremy Dreyer, founder and CEO of SkyHelm, and Logan Sarrington, one of their technical account managers. SkyHelm is an MSSP focused on utilities which mainly in the electric co-op space but with other utilities across transmission, distribution and generation, too. Their field perspective is exactly what this topic needed.

R1 Part 1.3: Evaluate Anomalies and Determine Further Action

Anomaly investigation and response falls under the CIP-015 R1 Part 1.3. The requirement reads simply enough, until you break it down. There are really two obligations packed into one sentence.  

The first is the method: implement one or more ways to evaluate an anomaly. The second is the one people tend to skate past: determine further action. You don’t just evaluate what you’re going to do; you have to reach an outcome. At audit time, that shows up directly in your evidence, where you must document not only your methods (policy, processes, playbooks) but also the actions you took in response to detections.  

The technical rationale for CIP-015 expressly ties this evaluation back to CIP-008, and that union is where a lot of the nuance lives, as we’ll discuss. Alignment between CIP-015 R1 Part 1.3 and CIP-008 is critical. If there is a miss alignment of critera or escalation process a compliance gap may exist.  

Evaluation Process

Developing a repeatable process for evaluating anomalies is key to a successful CIP-015 program. Process development requires understanding the team's capabilities, anomaly detection methods, and the data sources that can provide context for those evaluating anomalies. Defining the expected outcomes can be done, and the process can be developed from there. Procedures and playbooks should be maintained to support the process.  

Simple Evaluation Process

Anomaly Classification: Focus on the Action Required

In a properly designed OT network, genuine malicious activity is rare. Most detected anomalies are false positives: completely benign and untracked. For example, someone plugged into something at a substation they weren’t supposed to and didn’t tell the right people.  

Evaluating anomalies involves classifying them, but the classic true-positive/false-positive grid invites confusion because it mixes two different questions: did the alert fire versus was there malicious intent behind the alert firing? The latter is much more relevant.  

The CIP-015 drafting team wisely sidestepped this confusion by focusing not on labels but instead on the action to be taken when an anomaly is detected. The technical rationale lays out four actions plus a fifth “choose your own adventure” category. In practice, you must layer on several other common actions to address that grey area between malicious and non-malicious false positives.

The technical rationale actions (left) and what teams actually do in the field (right).

A good example of that grey area is an alert that says “unknown device connected to RTU.” It could be an incident, or it could be completely routine, depending on whether the right person plugged into the right machine and is doing something appropriate. Let’s drill down into the first two actions in the technical rationale and related common actions.

No Action –> Change Ticket Correlation

The first thing a SOC analyst should do when they see an anomaly alert is try to correlate it with a change ticket that explains the activity. Could it be tied to a failover event from the primary to the backup control center that radically changed traffic? Matching anomalous activity to a change ticket offers proof that no further action is warranted but only if the ticket has sufficient detail. When there’s no ticket or it’s too thin to correlate, that’s a signal to fix the upstream process, not just close the alert.

Further Investigation ­–> Acknowledgment into Baseline

It’s easy to think of your network baseline as something fairly static to be checked periodically. But baselines change over time. Software upgrades, new protocol variables, failovers and infrequent maintenance activity (such as annual relay or battery testing) impact your baseline. Still, you don’t want to add all of that activity into the baseline. Incorporating infrequent but normal activity into the network baseline effectively creates exceptions, leaving a wide berth for something malicious to use.

A better practice is to capture that traffic the first time it happens and document it in a knowledge base for fast correlation when the same maintenance operation recurs so the SOC is alerted when they happen and can identify what they are.

Beyond the Network Baseline: Recognizing “Normal” and Benign Deviations

For INSM, when you say baseline you’re referring to the network baseline. But to really understand what the system is doing, you need to expand the baseline concept past network flows. What is its purpose of the communication? For example, if an engineering system is reading voltages from metering and relays it will show up as polling via DNP3 or a vendor protocol. In itself its not possible to easily determine if this is benign or malicious. This is where context matters.

To the extent possible, having a single source of truth for your asset inventory, with uniform documentation, will accelerate investigations, among other things. A basic inventory that lists assets and their attributes has limited forensic value; you want full visibility into each device’s behavior and communications.

A good way to zero in on engineering anomalies in the expanded baseline is to pull the last year of control center change orders (usually a manageable number), group them, and pick 10 to 20 that the SOC team should recognize and be able to correlate readily when they happen again. Maybe the activities go into the baseline, maybe they trigger alerts but are understood as either normal or benign, as long as they’re accounted for. Here are several common examples of each.  

Seasonal/Load-driven Operations and the “Fog of War”

A vivid example of how benign deviations can be exploited is during storm operations. Utilities in areas subject to hurricanes, winter snowstorms and other seasonal events understand this all too well. When a storm hits, the backup control center comes online, and read-only logins can spike by half as a response center fills with people. You can’t call the control room mid-storm to validate anomalies; safety and reliability come first.  

Adversaries exploit that so-called fog of war, timing their attacks to large storm events.The SOC is going to see a lot of alerts related to anomalous behavior. Documenting these storm-related deviations can save everyone time and headaches.  

The CIP-015 ↔ CIP-008 Handoff  

As with many of the NERC Reliability Standards, CIP-015 is linked to another CIP standard, in this case CIP-008. CIP-015 M1 Part 1.3 requires “documentation of escalation process(es) that could include Cyber Security Incident response plan(s).” Clearly, the two standards must be aligned, especially around definitions that determine actions.  

Activating CIP-008 should be rare on a well-designed network, but it should be practiced. The standard triggers reporting requirements for two things: attempts to compromise and cybersecurity incidents. if you call something an “attempt to compromise” in your CIP-015 evaluation and don’t activate CIP-008, you’re going to have a reporting issue. Come audit time, you may also be glad you took the time to document negative decisions. When you determine something is not an attempt or incident, and why. An auditor pulling sampled events for a given time period may look for it.  

Anomaly Investigation and Response: Lessons Learned and Common Pitfalls

As our INSM webinar series progresses, several recurring themes have emerged. Here are some lessons learned and common pitfalls we’re seeing as utilities design their anomaly investigation and response plans:

  • As with all things NERC CIP, words and definitions matter. Get alignment across teams and standards, and take documentation seriously.
  • Tuning is expected: if your team is re-clearing the same false-positive alerts every week or taking days to correlate ultimately benign activity, they can’t focus on the real goal:  finding malicious activity and proactive response. That doesn’t necessarily mean disabling alerts but rather muting them, as we’ve advised in previous sessions.
  • Baselining needs to occur consistently, ideally monthly through a change advisory board. If you’re a SOC manager and you’re not seeing some certain issues regularly incorporated into the baseline (or otherwise documented), your team is working harder  than they should.  

The two biggest pitfalls we consistently see apply well beyond this topic and well beyond INSM. One, designing a program to meet requirements is an effort in itself, but when what’s written on paper isn’t what the team actually does, it’s not hard for an auditor to figure that out. This hazard occurs when team that drafted the policies and procedures did so in a silo, and their work is never adjusted to match day-to-day reality. Over time, the two things are so far apart that you have to start over, often with outside help.

The second major pitfall we continue to see is utilities of all sizes treating CIP-015 as a project, or even just a technology project. It’s a substantial program with equal parts people, process and technology. Most utilities have only to look back at the last time they took a project approach for a major undertaking to realize why it failed or underperformed.

This post only scratches the surface of what was covered in the webinar. Register for the full NERC CIP-015 webinar series and watch the Episode 7 replay for the full discussion, including a flow diagram you can use to build your evaluation process and some thoughts on how to get your SOC and control room talking. If you’d like to talk through how any of this applies to your own program, reach out. We’re here to support you.

No items found.