📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

CISM EXAM PREP - Domain 4A - Incident Management Readiness

Inside Cloud and Security1:36:04

Transcription

Welcome back to the CISM exam prep series. So, domain 4 is all about incident management. And in domain 4 part A, we're going to focus on incident management readiness, beginning with the incident response plan. Then we'll move on to other key elements of business continuity management that support our incident response efforts, including business impact analysis, the business continuity plan, the disaster recovery plan, as well as our incident classification and categorization strategy. And with our processes in order, we'll wrap with a look at incident management training, testing, and evaluation. This is where we connect our people to our incident management processes, ensuring the organization is ready to respond when the need [Music] arises.

In domain 4 part A, we'll be focused on incident management readiness. And while you may know me from my exam prep content here on YouTube, in my 9-to-5, I'm a cybersecurity strategist and a VCSO for a regional bank where I'm exercising my cybersecurity knowledge every day, like that which you'll find here on the CISM exam. More importantly, last year I helped thousands achieve cybersecurity certifications like the CISSP, CCSP, and the Security Plus exams. And I'm bringing that proven formula to you here with the CISM exam. As with all my exam prep courses, you'll find a PDF copy of this presentation available for you in the video description to leverage in your exam preparation as you require. You'll also find a clickable table of contents available in the video description, and it should appear automatically on your YouTube video timeline.

So we are beginning the fourth and final domain, which focuses on incident management, and in the series, we're in the first half of domain 4, 4A. So we'll have one more installment after this one. So if we look at the syllabus for domain 4 at a high level here, part A includes incident response plan, business impact analysis, the business continuity plan, disaster recovery plan, incident classification and categorization, and we'll wrap up with incident management, training, testing, and evaluation.

So if we look at the supporting tasks of those 37 and which map to domain 4A, we have number 29, establish and maintain an incident response plan in alignment with the business continuity plan and the disaster recovery plan. So remembering that incident response plans and the disaster recovery plan support the business continuity plan, which serves the business. Number 30, establish and maintain an information security incident classification and categorization process. This is what we're using to establish the incident type. So, what kind of incident is it? And the impact and severity. So, how bad is the incident? 31. Develop and implement processes to ensure the timely identification of information security incidents. 32. Establish and maintain processes to investigate and document information security incidents in accordance with legal and regulatory requirements. And then establish and maintain incident handling process, including containment, notification, escalation, eradication, and recovery. A lot of terms you definitely want to be familiar with there. You want to know at the end of the day, the incident response flow. And you want to be familiar with all of these terms. 34. Organize, train, equip, and assign responsibilities to incident response teams, ensuring teams have the skills they need, the resources they need, and everyone knows their role. 35. Establish and maintain incident communication plans and processes for internal and external parties. So we need to choose channel, verbosity, and the cadence of communication based on our audience. We're always tailoring that communication. And 36, evaluate incident management plan through testing and review, including tabletop exercises, checklist review, and simulation testing at planned intervals. So plans require regular maintenance and testing to ensure they remain effective.

So in section 4A1, we're focused on the incident response plan. We'll look at incident management versus incident response. We'll talk through our desired outcomes, how policies and standards factor, response and recovery plans, teams and staffing, and we'll wrap with a discussion of notification.

So what is incident management? Well, incident management is essentially part of an organization's risk management program. The focus is on proactive planning, preparation, and identifying risks and abnormal events. The goal is quite simply to minimize negative business impact and to return the business to normal operations as soon as possible.

So what are the core elements and the goals of incident management? Well, it aims to address a broad spectrum of potential disruption. Certainly, returning the business back to normal operations quickly, but let's go a bit deeper. So, for example, we want to address situations like losing a critical staff member, IT outages, market changes, major disasters. Our goals include minimizing harm, restoring normal operations quickly. Some key considerations here in developing an incident management procedure would include tailoring our process to the organization. There's no one-size-fits-all approach. Certainly, there are processes in terms of flow we can rely on as a template, but how we approach incident management will be influenced by the organization's operational maturity, its size, our industry, and our compliance needs. And this process definitely needs stakeholder buy-in. Especially senior management needs to be on board with our incident management program. And incident management addresses the full incident life cycle: before, during, and after an incident.

So let's compare incident management versus incident response in terms of goals. So stated simply, the primary goal of incident management and incident response together is to minimize the impact of security and incidents on the organization. So if we get a bit more specific, first and foremost is protecting human life and safety, minimizing downtime and business disruption, maintaining confidentiality, integrity, and availability – the three legs of the CIA triad, the foundation of security, preventing recurrence, complying with legal and regulatory requirements, protecting the organization's reputation, which can be difficult to destroy and that loss of reputation immeasurable. So, while there's not strictly a numeric order here in terms of priority, we can put a couple of numbers by this safely. Protecting human life and safety is always number one. Minimizing downtime and business disruption, very close second.

So let's dig into some key processes supporting our goals in terms of incident management and response. Early detection and rapid response. We want to be prepared to identify incidents quickly and contain the threats fast so we minimize their impact. Minimizing that adverse impact. So limiting the damage and maintaining and restoring business continuity, preserving evidence. So this will support investigations later, allow us to get down to root cause, address any legal needs should there be litigation, maintaining a chain of custody, very important, meeting our obligations from a legal and regulatory perspective, as well as any contract service level agreements. Protecting the organization's reputation and our stakeholders, maintaining trust, and keeping stakeholders informed appropriately throughout the incident, establishing clear roles and responsibilities. So, in the program, we're going to define who does what from a technical perspective, a management perspective, as well as the external players in the incident management and response process. Facilitating continuous improvement. So we always want to conduct post-incident reviews, learn what we can learn, and update our plans accordingly. Provide adequate funding and resources. This is where support from the top comes into play. Budget for tools, training, and personnel. If we don't have management support, that budget is likely to be lacking. And aligning with business objectives and risk appetite. So, tailoring our approach to the organization's risk tolerance and balancing cost versus asset value. But I still haven't really told you the difference between incident management and incident response. Stick with me a couple of minutes, and we'll get there.

So, let's talk through the process steps in the incident handling life cycle. So, here, for example, are the steps that roughly align to the SANS incident response framework. Preparation. So planning, policies, tools, training, management, support. Identification. So how we detect events and verify incidents. The initial analysis, sometimes called triage, and prioritization, where we apply the classification and categorization. Containment, limiting the scope and magnitude of the incident, preserving evidence, that chain of custody. Eradication, removing the root cause of the incident. Recovery, restoring affected systems and services to normal operation while attempting to meet our recovery time and recovery point objective. So, recovering quickly and restoring the data we need. So, you notice some of the key terms I said you should definitely be familiar with on exam day. So, you see here containment, eradication, and recovery. Remember these lessons learned. So post-incident, we analyze the incident, identify improvements, and update plans and procedures. So this is one flavor of the incident handling life cycle. So we could also look at it from the Carnegie Mellon University perspective. And here we see prepare, protect, so safeguarding assets. Detect, where we identify suspicious activities. Triage, where we sort, categorize, and prioritize those incoming events to then declare an incident and respond, take action, the technical actions to respond to the incident, and then the management and legal actions that may come alongside. So that's another perspective. We could also look at it from the NIST perspective, which includes preparation, detection and analysis. Then the containment, eradication, and recovery steps. And we really go through a process there of detecting potentially multiple aspects of an incident. And then there's the post-incident activity, where we find our lessons learned through that post-incident analysis and we update our process accordingly.

So, if I were to give it to you from just a general flow perspective, I think you'll tie it all together here. But NIST is covered in their Special Publication 800-61. You really need to focus on the elements of these processes and know containment, eradication, and recovery terminology, what the difference is between those three. I've covered it a couple of times here. But from a general flow perspective, if we just look at this from a general process, there's the planning and preparation, where we create our policies, we get support, we build awareness, we have the tools in place. Detection, triage, and investigation, where we define an event versus an incident, because every incident will include events and alerts, generally speaking, but not every event and alert is an incident. So we have to detect, validate, prioritize. Then we contain. So here we're executing containment, performing forensics, initiating that recovery process, finding the root cause. And then we have the post-incident assessment, where we evaluate our response. We identify opportunities for improvement, corrective actions. And we report our metrics. So how did we do? Did we recover meeting our recovery time and recovery point objectives? And then there is incident closure. The final analysis. We submit the reports and incorporate those lessons learned into our incident response process.

All right, you've been very patient. So now let's focus up on incident management versus incident response. What is the difference exactly? Incident management is the strategic program and the overarching framework. Incident response is the process that is triggered when an incident is declared. Incident management includes planning, policies, resource allocation, governance, recovery, continuous improvement. Incident response focuses on immediate action: identify, contain, eradicate, recover. Those terms you need to be very familiar with. Ensures preparedness and alignment with business goals. So, incident management maintains that higher-level, business-centric, big-picture focus. Incident response, the goal is to minimize the immediate impact of the incident and restore operations quickly. So incident management, think big-picture governance, and incident response, think boots-on-the-ground action. So the relationship here is incident response executes part of the overall incident management strategy. Together, they build resilience in service operation. And remember, senior management support is crucial for both. We have to have the support, the resources, and the budget. And remember the terms I said are important here: containment, eradication, and recovery. Make sure you know the difference between the three. You'll hear them again in the series, but there's one place where you can go back to refer at any time.

All right, moving on. Let's talk about the information security manager's role in incident management and recovery. So their involvement varies based on the organization's size, the structure, industry, regulation, organization's maturity from a program and process perspective, but we can identify some typical responsibilities. The information security manager often leads or participates in security incident response. They certainly assist with business impact analysis. They'll coordinate recovery solutions with IT. They act as an advisory, as a consultant to the organization. They have to understand business continuity, disaster recovery, and incident management requirements so they can ensure security plans align well with the broader continuity and recovery objectives of the organization. That security-business alignment we've talked about throughout the series.

So incident management and risk management are actually two sides of the same coin. They do differ in their focus. So risk management focuses on prevention, where we're putting our controls in place. Incident management focuses on effective response. So a materialized risk is an incident, and incident management establishes that plan for coordinated response to contain damage and prevent disaster. Incident response capabilities reinforce the overall risk posture. The goal at the end of the day is prevent escalation, contain the incident, and prevent it from becoming a disaster. But incident management, remember, establishes the plan. Incident response capabilities are the boots on the ground.

So, integrating assurance processes. We've talked about assurance processes a few times in the series. So this is the practice of systematically connecting and coordinating organizational functions like physical security, HR, and legal with our incident response plans. In this case, the key value of assurance process integration to the business is to ensure smoother and more effective collaboration during high-pressure incidents by clearly defining departmental roles, which is going to eliminate confusion under stressful conditions and unintentional overlaps during crisis, so everybody stays in their lane. Uh, testing interdepartmental linkages. So this proactively identifies and addresses potential communication or coordination breakdowns. Facilitates a unified response. This leads to a more efficient and effective incident resolution. So just elements of planning and validation to make sure that when the situation arises, the organization can respond competently and effectively.

So let's talk value delivery. So to deliver maximum value, an effective incident management program should seamlessly integrate with business processes and structures, minimizing disruption to the business and maximizing efficiency in a response. It strengthens the organization's risk management capabilities. Remember, risk management tends to be proactive. Incident management is the reactive that serves to support when our proactive measures don't do the job. Align with business continuity planning so that critical functions can continue running or be quickly restored. So supporting our business continuity plan, contribute to the enterprise's overall security strategy, and acting as a backstop to optimize risk management efforts of the organization. So the incident management program effectively serves as a final safeguard against threats, helping to contain damage and streamline recovery while continuously improving the risk posture of the organization.

So resource management in this context, including time, personnel, budget, really underpins the efficiency of any incident management effort. So effective triage is where it begins. Remember, triage is that critical initial process of rapidly assessing, prioritizing, and categorizing security alerts to determine which threats are most urgent and require our immediate attention. So, deploying our limited resources to the highest risk incidents first, protecting critical functions, enabling rapid resumption of service. Considerations here: resource identification, understanding our time constraints and the skills and availability of our personnel and our budget to apply to these processes. Triage, that's where we're assessing impact, urgency, and scope of incidents and prioritizing. Allocation, we prioritize critical incidents. We match skills to tasks, and we use our resources as efficiently as we can. And business continuity, focusing on core systems, maintaining communication with all relevant parties throughout the process, using the medium and language and cadence that best meets that audience.

So when we define incident management procedures, as you saw earlier, there is no universal approach to defining incident management procedures. Organizations adopt established frameworks. That might be NIST. It might be SANS. It might be that Carnegie Mellon. It might be something of their own design that simply uses these frameworks as a starting point. They're customizing the response, tailored to the organization. At the end of the day, the goal is simply to match best practices with the enterprise's environment and their risk profile. So for you on this exam, we've talked through several aspects of the flow of incident response. We've talked through the difference between incident management and response, and those key terms you need to know for exam day. That's where you should focus. Don't worry about rote memorization of any specific framework.

So let's talk about assessing the organization's capability. Answering that question, "Where do we stand?" by assessing current incident response capability to identify where the organization needs to change. There are a few techniques we can use here. You've heard some of these before. So something as simple as a survey, talking to our business leaders, managers, to the IT department, self-assessments versus best practices, external audits or external assessments, which would really be a thorough, unbiased evaluation. Uh, reviewing our incident history to identify trends, weaknesses, impacts, identifying those areas where we could have responded more effectively. The less frequent but catastrophic events like a large-scale earthquake might be addressed through specialized insurance or advanced DR planning. So those will be outliers, but still need to be addressed in some way. So performing a gap analysis compares current versus desired incident response capability. So a gap analysis clarifies which processes need improvement for efficiency and effectiveness, and any technology or skills required to meet those response goals. What steps to prioritize based on impact and cost-benefit considerations. The findings of a gap analysis can provide direction on development or update of the incident response plan.

So incident management plans also need to cover operational continuity logistics. So procedures, having detailed, accessible hard copy procedures in case electronic systems are down or staff is unavailable, which can definitely happen in a disaster situation. Off-site storage. So securing essential business forms and software off-site to enable alternate work locations if needed. And understanding dependencies, ensuring key hardware and software dependencies are documented, including backup availability. But all of these considerations ensure members of the incident response team have the resources and the information they need to respond competently and to minimize the need to make ad hoc decisions under stressful circumstances.

Then there's defining our response team roles. We need to make sure that clear roles and responsibilities for each of the teams is defined before an incident happens. So there are a number of teams that are commonly involved here. You have the emergency action team, which will deal with immediate threats, damage assessment, basically assessing the extent of damage. Emergency management, handling some coordination and decisions in disaster scenarios. The relocation team, helping folks move to alternate sites. The security team, who focuses on containment and forensics. But the teams should have clarity on decisions like service priorities, the required staffing, and the location of alternate facilities well before any event happens.

So then there is the training of our team. So ensuring that our team is well trained so their response is competent when the time comes. And the information security manager should conduct or coordinate scenario-based exercises to test response steps. Confirm resource adequacy, including tools, documentation, availability of staff. They should mentor new team members on processes, roles, responsibilities, and corporate standards to make sure everybody is prepared for game day. And that's where they can use on-the-job training for continuous improvement in policies, standards, and tools. So, we can get our folks well acquainted with our policies and our standards, and then they can help us in the process of continuously improving those. Scheduling formal training when necessary for team members to achieve certifications or competencies relevant to their role in the process.

Now, we've mentioned the importance of notification a couple of times, making sure that we have timely notification, which is going to be vital to reducing damage. So, an effective approach will include automated alerts, emails, SMS, etc. Documented escalation paths, a list of relevant stakeholders to make sure that we're reaching out and informing the right groups at the right time. This will include management, HR, legal, IT, security, a broad range across the business. We need to define and communicate responsibilities. So roles and responsibilities should happen well before the event. And we want to formalize those roles and responsibilities in job descriptions wherever we can. So it leaves no room for interpretation.

So let's talk challenges in developing the incident management plan. So developing the plan comes with challenges of its own. And the information security manager, your role on this exam should anticipate and address these challenges proactively. So common obstacles will include a lack of management buy-in. So without senior leadership support, resources may be limited, and responses may be slow or inconsistent. We know that a lack of management buy-in is a fatal flaw in terms of our odds at being successful in developing an effective incident management plan. Misalignment with organizational goals. So rapid change in the organization can sometimes outpace updates to incident plans, leaving new or expanding areas unprotected. So leave us with gaps. We might have turnover in the team. If we lose key champions, that can stall plan development and maintenance. Ineffective communication. So either too little communication, which can lead to confusion, too much communication, which can lead to stakeholder fatigue where people just begin to ignore the messages, or a failure to tailor communication to the audience, being too technical with our business leaders or too vague with our technical teams, for example. And overly complex plans. If a plan is too broad or too complicated, stakeholders might not fully commit to it, leaving us with gaps when it comes to execution.

All right, let's test what we've learned with a practice question. So when a possible breach of an organization's business system is reported by a SOC analyst, what should be the first action taken by the incident response manager? Is it A validate the incident? B run a vulnerability scan on the affected system? C disable the affected user's account? Or D review the system logs? So looking at what they're asking for here, what is the first action that should be taken when we have a potential breach reported? So we have a potential incident here. So what is the incident response manager's first step? Well, just looking at this, since they're asking us for the first action when a potential breach is reported, I'm not going to run a vulnerability scan because I don't know that I have an incident yet. And that's really just going to reveal a vulnerability, right? It's not going to detect a breach. Disabling the affected user's account, potentially impacting a process, a business process, not going to be my first step. So, I might review system logs. That would certainly allow me to potentially detect whether or not that breach is real. Or do I validate the incident? So when investigating an incident, it should be validated. As we've discussed in the past, every incident will include alerts, but not every alert is an incident. So when we're investigating a possible incident, first we want to validate. We want to confirm it's an incident. This might require some initial triage to confirm before we declare that we have an incident on our hands.

And that brings us to section 4A2, where we're going to focus on the goals, elements, information collected, and benefits of the business impact analysis. Essentially, we're going to distill the key aspects of the business impact analysis you should know for exam day.

So what is the BIA? Well, its purpose is to determine the impact on the business if critical processes or resources fail. It's going to be the foundation for several other documents and plans. It's the foundation for risk management, our business continuity planning, and the disaster recovery plan, which lives within the business continuity plan. But the business impact analysis helps the business know where to focus in all of these. So, it's going to answer questions like: How fast does a disruption hurt us? What do we need to recover? What's the overall operational hit?

So if we look at it in terms of the three security dimensions, those in the CIA triad, we have a loss of confidentiality. So unauthorized data disclosure, which can lead to reputation damage, fines, legal issues. Loss of integrity. So unauthorized data or system changes, which can lead to data fraud or system compromise. A loss of availability. So critical systems or processes being down. This leads to loss productivity, revenue loss, and certainly customer dissatisfaction. So the three core goals of the business impact analysis: It needs to consider how disruptions can affect each dimension of information security: confidentiality, integrity, and availability, effectively the CIA triad.

Now, the BIA also focuses on three overarching objectives which guide organizations in prioritizing tasks and resources during incidents. So there's prioritizing criticality, where we identify and rank business functions by potential disruption impact. This focuses on recovery of what matters most first. Downtime estimation. So determining maximum tolerable downtime for processes. This defines the acceptable interruption window and resource requirements. Identifying essential people, technology, facilities, etc., needed for recovery. This clarifies what is needed to get back online and running again.

So key considerations during our business impact analysis. An effective approach is to analyze a manageable set of realistic scenarios with a broader group of key stakeholders. So we want to use realistic scenarios. Don't only plan for absolute worst-case. Consider likely disruptions, which may be lesser in their severity. We want to avoid inflating the perception of risk. We want to understand loss escalation. So how impacts worsen over time. This helps us to set realistic recovery time and recovery point objectives. Assessing tangible and intangible impact. So measuring financial loss and evaluating intangibles like damage to the organization's reputation, aligning with other risk analyses. So ensuring that our BIA findings integrate with broader business risk assessments. So a common pitfall we see in business impact analysis is too much focus on worst-case outcomes, on outliers, which can inflate perceived risks. It can have us focusing in places other than where we need to be and potentially overspending in the process. So go back to the first point on this slide here: Use realistic scenarios. Plan for likely disruptions, not just the worst case.

So common business impact analysis steps. So here are some core considerations and steps that we see in most approaches. We identify business functions. We detail the activities within each function. We map internal and external dependencies, so we know everything that's needed to get that process back online. We sequence the critical operations. So once we understand the dependencies, then we can understand in what order we need to get them back online. Establish recovery time objectives. So how fast must it be recovered? Define resource needs for restoration. Determine manual workarounds and our recovery point objectives. So how much data loss is tolerable. Assess maximum tolerable downtime. Analyze upstream and downstream impacts on other functions. So you're looking at some dependencies there. Compile findings to inform the business continuity plan.

So let's talk about information typically gathered in a business impact analysis. So we're looking for function descriptions and objectives, dependencies, impacts – financial, operational, customer, competitive, etc. So who is affected? How significantly? Work backlog accumulation rate. So we're now a bit more business function focused. Let's step into the technology side. So technology needs: servers, network, software, restoration procedures, and alternate site capabilities, manual workaround feasibility and duration. Remote work capability. We see this commonly when a primary site goes down. We have a plan for remote work so the people can continue to support business functions. Critical resource identification in terms of people and equipment. Potential alternate solutions. That's where we get creative and we come up with temporary workarounds where possible. Recovery sequencing and timeline. So, order of operations and time that it's going to take to bring those functions back online. But you've got a bit more technical detail down there in some of those questions. So, really, how and how fast can we recover? What recovery options do we have? And how long will recovery take? That's really what we're digging into in this longer list here if I wanted to distill it down. How and how fast can we recover? What options do we have? And how long is it going to take?

So if we look at that business continuity management picture again, the business impact analysis is going to feed some of our key metrics like recovery point objective, recovery time objective, the max maximum tolerable downtime. That's going to feed into our business continuity plan and our disaster recovery plan. The technical aspects lie within the BCP, and then over time, we need to test and maintain the plans to make sure that they work and remain current.

So what are the benefits of a business impact analysis? Well, it helps us understand potential losses. It quantifies financial and operational impact. It helps us to prioritize recovery. It focuses our efforts on the most critical functions. Helps us identify dependencies, revealing critical links between processes and resources necessary to bring our critical business functions back online. Helps us to improve incident response awareness. It educates the organization on impacts and our recovery needs. The BIA tells you what's critical, how bad it is if it breaks, how long do we have to fix it, and what we need to do it. But the BIA essentially drives an intelligent business continuity and disaster recovery strategy.

All right, let's test what we've learned with a practice question. What are the most essential elements or components outputs of a business impact analysis? The most essential outputs of a business impact analysis. Is it A, business continuity testing methodology that we're deploying? Is it B, the cost of business outages in a year as a percentage of our security budget? Is it C, our downtime tolerance, our resources, and criticality? Or is it D, the structure of the crisis management team? So, there are a couple of things we can knock off here right away. So, the business impact analysis doesn't get into the weeds on business continuity testing methodologies. We didn't even talk about that. Now, it does help us to quantify different metrics like the financial impact, but we're not going to extrapolate that into the cost of business outages in a year as a percentage of security budget. That's going beyond our focus of BIA considerably. So we have downtime tolerance, resources, and criticality, or structure of the crisis management team. So the most essential output here would be C, which is downtime tolerance, resource list, and or criticality of those resources. So the main purpose of the BIA is to measure the downtime tolerance, the associated resources, and the criticality of a business function. So options A, B, and D are all associated with business continuity planning, but they're not part of the BIA, which feeds into the business continuity plan.

All right. And that brings us to section 4A3, which is the next logical step, the business continuity plan. So here we'll talk about integration with incident response and business continuity, methods for providing continuity, high availability considerations, defining a risk management framework, uh, insurance and cyber insurance considerations, and we're going to touch on some technology concepts here when we're talking about high availability, but as in other areas of the series, we're going to keep it appropriately CISM level, which is to say, not too technical.

So what is a business continuity plan? So defined, it's the plan that outlines how the business continues delivering products and services during unexpected interruptions. Core objectives: minimize overall impact, protect critical processes, make sure we define roles. Key considerations: essential processes and personnel, necessary IT assets, critical third-party dependencies, and we're going to integrate business continuity and the BCP with incident response and disaster recovery because business continuity planning has to be integrated with those two. In fact, disaster recovery is a component of the business continuity plan. So they are all connected. Incident response handles the immediate incident. Disaster recovery restores it after major disruption. Business continuity keeps business functions going. We need to align the plan with our recovery time and recovery point objectives, which come from the business impact analysis. They have to be consistent across all three of these plans. Recovery is only as fast as the slowest of these. Bear that in mind. Consistent planning ensures smooth transitions in execution. So for example, declaring a disaster at the right time, moving to alternate sites, etc. So having consistent execution because we have a plan where we have documented these processes. Incident classification determines which plans get activated. So the classification is the type of incident. So if we go back to that event that happens, that business impact, the recovery point objective focuses on our data loss limit, which will direct how we back up our data, how we build that redundancy or recovery capability in. And the recovery time objective focuses on the length of the service outage, that's going to direct our high availability solutions we'll use for our services. How we'll build high availability, redundancy, or resiliency into the services.

So if we think about network continuity, so preventing single points of network failure from a redundancy perspective, building in extra capacity, backup components, multiple devices, multiple network paths, diverse routing, so alternate physical or logical routes immune to the same failure. For example, in an organization, we'll often see two internet service providers connecting that organization to the internet with redundancy in the internet routers that connect that business to those redundant providers. Alternate routing, so different carriers or routes, often with automatic switching. That's really what I'm describing there. And then last mile protection. So securing the final network link into your facility. So dual entry point. So if I were to describe that in the context of a big cloud service provider like an AWS or an Azure, they will have redundant internet connectivity points coming in, perhaps at different ends of the facility. So not only different internet service providers but coming into different physical points of ingress into that building. So the same issue doesn't necessarily affect both of those links.

So high availability, so keeping our critical systems accessible with minimal downtime, addressing the RTO need, redundant hardware and software, so clustering, load balancing, data replication, rapid failover, generally automatic switchover to backup systems where we can facilitate automatic switchover, though the impact is minimal or none at all. Geographic distribution, so spreading infrastructure to mitigate local disasters. And for organizations that use the cloud as their backup data center, this is more or less built in automatically unless you make a conscious decision not to. And if you're fully a public cloud customer, what you'll find is when the CSP chooses a data center pair, they will generally make sure there's at least 300 miles between those pairs. So for example, in Microsoft Azure, if US East on the east coast is your primary data center, the backup data center pairs on the west coast. So a good distance between them. Network reliability. So the redundant network devices, switches, routers, and connections.

So a few key terms and concepts you should be familiar with. Here we're we're dipping our toe into a bit of technology-related terminology as it comes to high availability. There's single point of failure. So any component whose failure halts the entire system, we want to identify and eliminate these if we can. System resilience. So resilience speaks to the ability to maintain acceptable performance during or after a disruption. High availability, so continuous operation with minimal downtime via redundancy or failover. And fault tolerant, so a system operating correctly even if components fail. This is really common in a redundant array of independent disks. We see this in hard drives in a server when you have a RAID array. If one physical disk in the array fails, the system continues to operate at near pre-failure capacity. Examples of these technologies, just a few examples that protect availability. You've got server redundancy. So clustering is pretty common for databases. Load balancing, very common for web servers, where we'll have a load balancing appliance of some sort that will distribute requests across an array of web servers. Network redundancy, we'll see alternate paths, diverse carrier entry points. So redundant internet service providers, redundant internet router connectivity, and coming in through dual entry points where we can power redundancy. So UPS, backup generator. So UPS accommodates graceful shutdown of a server that's uninterrupted power supply. Then the backup generator is intended for long-term operation. You'll see a diesel generator, for example, used to potentially keep a facility up for hours, days, or weeks. Failure management. So monitoring, predictive maintenance, and automated recovery is definitely preferred over manual where we can swing it.

Now for the expensive, rare events, that's often where insurance comes in. So the incident response plan should detail the enterprise's insurance coverage as part of overall risk management. So the role of insurance is risk transfer. It reduces financial impact. It does not prevent incidents, simply reduces the financial impact if that rare, expensive event happens. So from an integration perspective, we detail the coverage in the incident response plan, the limitations. So, understanding exclusions like regulatory non-compliance that might lead to a non-payment for a claim, or any deductibles that are involved. Key coverage types include business interruption, IT equipment, media reconstruction, the extra expenses that come with trying to recover after a significant incident. Uh, various types of cybersecurity coverage. We'll talk about cyber insurance in a moment. From a liability perspective, errors and omissions, we'll see that with where consultants are involved. And then fidelity, which covers dishonest or fraudulent behavior uh by an employee.

So let's dig into cyber insurance, which definitely can play a critical role in an organization's overall strategy. A common use is really just for transferring risk for a rare, high-cost, high-impact cyber event like a ransomware attack. We'll buy this for compliance reasons or because we, in our risk assessment, deem this necessary based on current forecast, past incidents, or just overall risk appetite. The critical point here is we have to understand the policy terms, conditions, exclusions, and the prerequisites for coverage. For example, in one cyber insurance policy I saw, the customer is required to call the insurer when they believe they have been impacted by a cyber incident, and during that incident management and response process, the insurer will be the central point for communication. They will drive notification as it needs to be driven. Key questions: ransomware payout rules. We always see the recommendations not to pay the ransomware attacker, but we'll need to see what under what circumstances the insurer will pay out for ransomware. Third-party breach coverage. So if we have someone in our supply chain, a vendor who's impacted, any type of proactive support requirements, any sort of controls that we need to have in place ahead of time, uh, incident reporting requirements, you know, as in that scenario I mentioned where the insured was required to call the insurer in the event of an incident, and the insurer would provide someone to drive management of the incident. But insurance carriers proactively identify risk factors generally for breaches, and they often require organizations to implement safeguards that reduce the likelihood or severity of incidents. So if you're being insured against ransomware attacks, they're going to make sure you have certain controls in place. The insurer mandates often require specific security controls, and leadership has a duty to review the policy requirements and ensure compliance to maintain coverage, to ensure that when the organization experiences one of these incidents and there is a breach, that the insurer will pay out because the org has followed the rules. But strong security controls reduce incident risk and support insurability, ensuring that our policy premiums are manageable and that we will get paid as an organization when the incident tells us we should. So for the exam, remember insurance is risk transfer, not elimination, and it's just one component of a broader risk management strategy. So alternate approaches include risk avoidance, reduction, acceptance; transfer is just one of those in the list.

So let's put what we've learned to use in a practice question. Which of the following is the most important element to ensure the successful recovery of a business during a disaster? Is it A, network redundancy maintained through separate providers? Definitely a good idea. B, hot site equipment needs are recertified on a regular basis. Also a good idea. C, appropriate declaration criteria have been established? Or is it D, detailed technical recovery plans are maintained offsite? So which of these is the most important to ensure the successful recovery of the business during a disaster? There are a couple we can knock off the list right away. So network redundancy maintained through separate providers. That is certainly one element we can put in place, but it's not necessary. Maybe we have another plan with a different facility rather than network redundancy at our primary site. Hot site equipment needs are recertified on a regular basis. What if we're not using a hot site? We can also use a cold site or a warm site. And that's really a bit more in the weeds. That's more of a disaster recovery item, not a business continuity item. Uh, C, appropriate declaration criteria have been established? Or D, detailed technical recovery plans are maintained offsite. Both good ideas. There's a clear winner here, and it is D. Detailed technical recovery plans are maintained offsite. So in a major disaster, the staff can be injured or prevented from reaching the hot site, resulting in the loss of technical skills and business knowledge. Therefore, we must maintain updated copies of our recovery plans at an off-site location. So, even if we're imperfect in our declaration of the incident, we can still recover if we have the plans. Even if we have people missing, we can still recover if we have the plans, the steps to follow to get us back in operation.

All right, that brings us to section 4A4, which is the disaster recovery plan. So we'll talk about business continuity and disaster recovery procedures, recovery operations, evaluating recovery strategies, addressing potential threats, and recovery site strategies. So we'll get into the different types of recovery sites and talk a bit about site selection, and finally, response and recovery strategies.

So what is a disaster recovery plan? Well, the focus of the DRP is on restoring IT systems, services, and data after a disaster. Whether that's failure, attacks, natural events, the purpose is to get critical technology operational again to support business recovery. So you see there's an obvious technical flint. That's why the disaster recovery plan lives within the business continuity plan. It is a component of the business continuity effort, and the phases in the DR process will vary across standards, frameworks, and technical exams. I'm trying to guide you down the language we see from ISACA for the CISM exam.

So how do we build and manage the disaster recovery process? Well, the phases of the DR process we see cited by ISACA include conducting a risk assessment and business impact analysis, defining a response and recovery strategy, documenting DR and business recovery plans, training our teams in these procedures, and then updating and testing the plans regularly. Finally, auditing the plans to ensure they remain appropriate and effective over time.

So how are the disaster recovery plan and the business continuity plan related? Well, the DRP activates after a disaster occurs. BCP includes prevention, mitigation, and disaster recovery. Key inputs: resource availability, required service levels, threats, and risk appetite. These inputs are going to guide appropriate response. Our evaluation factors include balancing recovery speed, the recovery time objective, against costs, preparation, activation, and considering insurance and related risks. But it's a balance between speed and cost. The DR focus is technical. The business continuity focus is the business. It's the broader business focus. So once a temporary or alternate DR site is operational, the business continuity team monitors the progress of restoring the primary site. They also conduct tests to determine if the site is safe to switch back. And when returning to the primary site, typically the same team that set up the DR site will coordinate and test the main data center. And after confirming it can operate normally, they transition operations back. Now, if the primary site is destroyed, the organization may move to a permanent alternate site or rebuild the original site depending on the long-term viability and the cost considerations. They'll have to weigh the circumstances as they encounter them.

Security considerations. So even during the chaos of a disaster, information security must remain in force to protect systems and data. We have to maintain CIA. The focus of the site is to get infrastructure and other technology online to support the business, which is why it fits into the disaster recovery plan because it's supporting the technology. The goal is to provide an alternate location for IT processing if the primary site fails. That is by definition the recovery site. We'll have several factors that drive selection. So business needs, which we'll pick up in the business impact analysis. The recovery time and recovery point objectives that come as outputs of that BIA, our budget, the organization's risk profile. We do need to remember geographic separation, ensuring the recovery site won't be hit by the same disaster. If we have a recovery site that's located across town and we live in a coastal city that's hit by a hurricane, or we live in a tornado-prone district region, then we can potentially have a primary and a recovery site both impacted by the same disaster. That's why we need to put significant geographic distance between them. Rule of thumb, as I've seen it in the field, is 300 miles minimum. As I mentioned, in the public cloud, we often see much greater distances than that. Other considerations: your compliance requirements and proximity to other hazards.

So, if we were to look at the types of recovery sites, there are three primary types: the hot, warm, and cold sites. So, let's break these down, and then we'll go into some alternatives to these. So a cold site is effectively just data center space, power, and network connectivity that's ready and waiting whenever you might need it. So to recover in a cold site scenario, your engineering and logistical support teams can move your hardware into the data center and get you back up and running with considerable effort because there's nothing plugged in and running or restored there. So the cost of a cold site is at the low end of the scale, but the recovery effort is going to be at the high end of the scale. Now, the warm site is a preventative site that allows you to pre-install your hardware and preconfigure your bandwidth. And when a disaster strikes, all you have to do is load your software and data and restore your business system. So the cost and effort are going to be somewhere in the middle between a cold and a hot site. The hot site allows you to keep your servers and a live backup site running in the event of a disaster. You're essentially replicating your production environment in that backup data center. So to recover allows for a near immediate cutover in

case of disaster at your primary site. So a hot site is a must for mission critical sites and services. However, the cost of maintaining that hot site is going to be high. The recovery effort, of course, is lower.

So, other recovery models. There are a number of ways that sites can be implemented. There's a mirror site, which is an identical active site operating concurrently, a type of hot site, effectively. A mobile site, which is self-contained, relocatable units like trailers, often duplicate sites, or recovery sites that are designed to be functionally similar or identical to the primary site. And that can range from a standby hot site, you know, fully operational and ready to go, to just facilities available through a reciprocal agreement with another company.

There's disaster recovery as a service, where we use a cloud provider to handle replication and recovery. That means we're using the public cloud like Amazon's AWS, Microsoft Azure, Google Cloud Platform as our recovery site, as our backup data center. This offers geographic diversity. It's often operationally expense-based. So, we're paying as we go, paying for what we use, means we're out of the data center business in that case.

There's the mutual assistance agreement, which is an agreement with other organizations for a cooperative disaster response. It's low cost, but it's high risk because both organizations might be impacted by the same disaster. We could have confidentiality concerns. There are certainly enforceability concerns, especially if the other organization is impacted. It's uncommon to see a mutual assistance agreement due to those difficulties in enforcement. The recommendation is to just avoid these if at all possible.

So, key metrics driving site selection: recovery time objective, so how quickly must it be restored? Recovery point objective, so how much data loss is acceptable? So then the acceptable interruption window, which is max tolerable business process downtime, which is related to the max tolerable downtime, the absolute longest the business can survive disruption. And then the service delivery objective, the minimum service level needed during recovery. So, what level of service are we going to commit to providing during recovery conditions? You want to be familiar with all five of these terms and acronyms here.

Key takeaway, though, is with shorter RTO and RPO, a faster, more expensive recovery site is going to be needed. So, if we have a short RTO, RPO, we need a hot site, a mirror site, disaster recovery as a service, one of those. And again, just looking at the picture, business impact analysis feeds into those metrics we discussed. Most notably, you're going to see RPO and RTO. Those are the most likely to appear on the exam. You might see max tolerable downtime or the acceptable interruption window.

Remember, business continuity plan focuses on the larger business continuity effort, and the disaster recovery plan addresses the technology needs that support the business. Within and over time, we test and maintain those plans to ensure they remain effective. So, I hope this visual here kind of helps you to piece together all of the many terms and acronyms that you encounter in the BCP, DRP, and incident response discussion.

So, let's talk about the disaster recovery plan essentials. It needs to be written in clear, simple language that everyone can understand, all of the players who are working with it. It needs to be easily accessible. We need to keep copies offsite so, in case we have an impact and some technology is unavailable, we have a copy in our hands that we can leverage to execute the disaster recovery effort. It should include activation trigger. So, when do we declare a disaster? Everyone's roles and responsibilities, ideally tied to job titles as well. Contact list. So, internal and external, and potentially how we'll contact customers in this case. Step-by-step procedures, making sure that folks know exactly what they need to do, the steps they will execute to minimize the need for ad hoc decision-making under stressful circumstances. Resource inventory, all of the the components we're going to be working with in the recovery effort. Evacuation plans in terms of protecting human safety if we have an impact at the primary site. And our communication strategy, which will include who we're contacting, the medium we will use to contact them, and of course, we need to tailor the language and the cadence, which should be steering us from that plan. Remember, DRP focuses on technology, BCP focuses on business. That should be second nature to you at this point in time.

So, let's talk about backup strategies. So, if we look at the three fundamental backup strategies, you have a full backup, which backs up everything selected, all the data. Simple restore, you basically restore the full backup. There's an incremental backup, which backs up all the changes since the last backup of any type. So, it's a fast backup because you're just getting changes since the last full backup or the last prior backup, but it's a slower, complex restore because an incremental backup is getting all of the changes since the last incremental backup. So, to restore, you have to restore a full backup and then each of the incrementals you've taken since that time up to your recovery point objective. A differential backup, on the other hand, backs up all changes since the last full backup. So, that's going to be a faster restore than incremental and a bit simpler because you're going to restore the last full backup and then the last differential, which gets you everything since the last full. And we need to consider scope, frequency, the media we're using. So, disk-to-disk backup or disk-to-cloud is strongly preferred. We don't see tape backup much anymore. You need to test your backups frequently to make sure they function. That is the most valid way to make sure that a backup is good, that you can actually restore it and the data is good and the system is functional and up to that point you have restored.

So, some other basics here. Let's talk about some backup strategies and their relative recovery speeds from faster to slower. So, electronic vaulting is where we're transferring backup data, often in bulk or batch, to off-site storage. Uh, remote journaling, transmitting transaction logs or data changes continuously offsite for point-in-time recovery. Or remote mirroring, where we maintain an exact data replica at a remote site. So, that could be asynchronous. So, there could be some latency there. Electronic vaulting, we'll see where we're going, you know, directly from server to cloud is going to be a great way to make that happen. So, if we look at those three technologies here, they lie in that uh disaster recovery RPO, RTO diagram we looked at earlier.

Let's put what you've learned to the test with a practice question. So, which of the following would be the most important consideration for an organization defining its business continuity plan or disaster recovery plan? Is it a) setting up a backup site? b) the data backup frequency? c) aligning with recovery time objectives? or d) maintaining redundant systems? So, which of these is the most important? They're all important in some respect, but which is the most important? So, we can take data backup frequency right off the list because our data backup frequency needs may vary by the scenario. Our RTO and RPO will dictate how we're going to facilitate recovery. So, maintaining redundant systems, that's down in the weeds again and may vary based on our recovery time and point objectives. So, that leaves us with setting up a backup site, which is definitely important when we think about a primary site failure. We're aligning with recovery time objectives. So, one of these needs to come before the other. And that's really the key to the answer here and really a theme you're going to see on the CISM exam. The answer here is aligning with recovery time objectives. So, business continuity planning or or or DRP should align with the RTO. So, once we define the RTO, that's going to inform our strategy. And remember, the RTO and RPO come out of the business impact analysis, a key output that then feeds our BCP and DRP. So, when prioritizing systems for recovery efforts, the RTO is considered to ensure that the business recovers the most critical systems first.

That brings us to section 4 A5, incident classification and categorization. So, we'll talk about the difference between classification and categorization, as well as the escalation process for effective incident management and preparing our help desk and service desk for security incident identification and the processes we need to lay out to support them.

So, let's talk about categorization versus classification. So, categorization describes what kind of incident it is. The purpose here is to group incidents by their nature. For example, malware incidents, network outage, hardware failure. So, we're grouping by type. That helps us to route incidents to the right team and it enables reporting and trend analysis as well. So, we can see where we're seeing the most incidents and where folks are spending their time. The outcome is a standard way to understand types of incidents that are happening in our environment.

And then there is classification. Classification defines how bad the incident is. So, here we assess severity, business impact, and urgency. And classification determines the priority level, if it's critical, high, medium, low, etc. But the classification will guide response speed, resource allocation, and our escalation trigger. The outcome here is a clear priority, ensuring focus on the most significant business impacts first. So, again, categorization is the type, it's what kind, and classification is the priority or the impact.

So, escalation defines when, how, and to whom incidents are reported as severity increases or time passes. So, key elements of an escalation process include triggers. We need to establish triggers which define what events mandate escalation. Even seemingly routine issues like data corruption might signal a security event if it's sensitive, critical data. It's going to be defined by security management and leadership. We'll want to document authority here. We need to clearly state who approves specific actions and recovery steps and disaster declaration. All need to be laid out in documentation. But there needs to be a clear statement of authority so we know who has the right. The authority to escalate. Need to define roles, backups, time estimates. So, if somebody's out of the office, for example, we have a backup person in that role to handle escalation. And execution steps should follow the documented procedure. So, we should have an ex escalation path that goes all the way to the top. So, if we begin at tier one support, maybe it ends up at tier three support, maybe it ends up with a business process owner, but we want to have that process documented in writing. And if a step fails or time runs out, we escalate immediately to the next level.

And managing alert conditions. So, incident severity can increase over time if it's unresolved, triggering broader notifications. For example, if we have a service level agreement with a B2B customer, a passage of time may trigger the need to escalate so we can resolve that issue because there may be penalties when we breach our SLA. We also need to define notification recipients as part of our escalation process. So, we'll pre-list stakeholders often, incident response team, management, legal, HR, potentially public relations. There may be key partners or even customers, especially in B2B scenarios, but we'll allow some flexibility there so we can respond based on conditions. We use secure communication methods whenever possible. And there needs to be a decision point to to end an emergency, for example. So, once contained, someone to assess the damage, decide if it's a disaster or if normal operations can resume. And we coordinate communications at that point with PR and legal or the appropriate roles. And then post-escalation tasks, actions like activating disaster recovery, isolating threats, recovering data, testing systems, all guided by the recovery time objective. The goal here is to ensure that whatever comes, we have timely, appropriate response and bridge detection with full incident response or disaster response activation.

So, the service desk is often the first line of contact. They're fielding the first report of anomalies or user issues that indicate a potential security incident. So, we need to make sure that our service desk is trained to distinguish standard IT problems from security red flags. They need clear guidelines from security management. And there's going to be some training required there, including social engineering awareness to prevent manipulation. So, these are more generalists in the IT sense. So, they're going to need a bit more guidance than security team members would. Procedure is going to be very important that we must follow defined processes for identifying potential incidents and escalating immediately to the correct security or incident response team. So, we need to provide a lot of guidance for our service desk here. They're less skilled. They're less specialized. So, they're going to need clear step-by-step process guidance wherever we can give it to them.

So, for optimal incident handling, remember these viewpoints. So, pre-planned categorization, classification, and escalation processes minimize chaos and impact during an incident. Prioritization ensures focus on the most critical issues so we can optimize our resource utilization and containment speed. So, trained help desk staff and clear escalation paths enable faster containment. So, we train our staff. We have a clear path for escalation. We're going to get the incident under control more quickly.

All right, let's just test our knowledge here with a quick practice question. So, when developing incident response procedures involving servers hosting critical applications, who should be notified first? Do you think the business process owner, the operations manager, the information security manager, or the chief information officer? So, when we're developing that incident response procedure involving servers hosting critical applications, who should be notified first? So, we can take business process owner off the list there. So, when we're dealing with the technology, the business process side of the house is not going to be the first contact. Now, if we discover that there is a business process issue, uh, separate from technology, that'd be another story. But we're talking about servers. So, we can take operations manager off the list as well. We have the information security manager and the chief information officer. Folks that will fall into the more technical or security-focused side of the house. And there's a clear winner here, and that is the information security manager. So, that escalation process in critical situations would involve the information security manager as the first contact. Remember from our early discussions of the information security manager role. They serve as a bridge between technology and leadership and our business process owners. So, they'd be our first contact. So, appropriate escalation steps are invoked as necessary.

All right, that brings us to our final section in 4A, that is 4 A6, incident management training, testing, and evaluation. So, here we'll touch on metrics and indicators, performance measurements, updating recovery plans, testing, incident response, business continuity and DR plans, periodic testing of plans, test types and results, and incident management metrics and indicators.

So, incident management, training, testing, and evaluation ensure that the organization's incident response plan is effective. Personnel are prepared to handle security incidents. We need to define roles and responsibilities clearly during the planning phase, ideally attaching those to specific job roles. So, when we have a security incident, we know who the incident manager is, who's managing any forensic investigation, etc. Now, from a training perspective, we need to train all relevant personnel, and that would include security, but also the broader IT department, the help desk, management. And what they're trained on will depend on the role, but they need to be prepared to recognize, report, and handle incidents to understand the full incident life cycle. With management, it may be simply to help us in making that escalation call. From a testing perspective, validating plans, roles, responsibilities, and communication. And there are several testing methodologies: tabletop exercises, functional drills, full-scale tests. We'll dig into these in just a moment. Then evaluation. So, after every test or real incident, we need to evaluate our process and how we did, capture any lessons learned, find our strengths and weaknesses, and drive continuous improvement. So, that post-incident evaluation activity is super important to drive improvement.

Now, in terms of metrics and indicators with incident management, we have a specific purpose here. We're looking for metrics to help us measure effectiveness and efficiency of our response, to inform management how we're doing, to justify resources, to identify areas of improvement. So, key performance indicators need to be quantifiable activity measures like time to resolve, for example, tracking progress against specific targets or SLAs's. Key goal indicators, broader strategic measures, essentially like reduce incident count by X percentage, so we can assess achievement of high-level goals through those key goal indicators. So, common incident management metrics, there are quite a few, for example, number of reported incidents, number of detected incidents, average time to respond, average time to resolve. So, response time and resolution time, both measurable and actionable if we're not responding or resolving quickly enough. Total incidents resolved successfully. Incidents not resolved successfully. Days without an incident. Proactive or preventive measures implemented. So, we relate time-based metrics to recovery time objectives as a measure of recovery success. So, for example, average time to resolve, we could tie that back to our RTO.

So, the value of performance metrics is that the incident management team can self-assess, identify trends, align resources if we see deficiencies, and ensure we're meeting requirements. From a performance perspective, we use KPIs and the key goal indicators to track operational and strategic success. Why do we measure? Well, to demonstrate success. So, demonstrating success justifies the resources we're putting towards our incident management and response efforts. It builds confidence in the team when they see improvements in their metrics over time. It also allows us to identify opportunities for improvement, driving continuous improvement. It shows our response capability. We can track speed and accuracy trends over time and compare that versus our RTO. But metrics reveal trends over time and help in aligning resources or adjusting processes to meet business and regulatory requirements. So, if we're not up to snuff, that may be an indication that we need to adjust our process or potentially bring additional resources to bear in the situation.

So, why update recovery plans? Well, business strategies, applications, threats, and environments constantly change. Triggers for updating recovery plans might include strategy shifts, new applications or technologies, business changes including mergers and acquisitions, or new regulatory requirements, updates to our hardware and software, physical or environmental changes, new sites, new data centers, and there are maintenance activities involved here as well. So, scheduled reviews when we're looking at these plans to make sure our recovery plans are up to date and remain effective, revisions for major changes, testing where we can incorporate lessons learned. So, when we test periodically, we might find something small has changed and we need to update a plan. We might find something major has fallen through a gap due to a change in infrastructure or application or process, and we have a major change. Training. So, after updates and tests, there may be some training necessary to bring everybody up to speed on the new aspects of a process or environment. And general recordkeeping. So, changes, so change management, contracts with vendors that are involved in a recovery process, or general procedural changes. But the goal here is to maintain a current, aligned plan for effective restoration that we know will work when we need it to work.

So, testing our incident response and business continuity or disaster recovery plans. There are tests that should be conducted, and the information security manager plays a key role here in identifying critical business applications and the infrastructure required to support them, and then developing a prioritized testing strategy. Not necessarily done alone, certainly will be a collaborative effort, but the information security manager plays an important role here. So, we need to test all aspects regularly. The primary goal here is to validate readiness, both personnel and processes. Key objectives in our testing include identifying gaps, whether they're procedural gaps, roles where maybe we've left a person out, or we identify risks that we weren't aware of or hadn't considered before. Verifying assumptions like resources and dependencies that we recognize being true today and ensuring that we've accounted for all the resources we need to recover and any dependencies that might exist. Evaluating our strategies, are they still effective? Do the plans still work? Checking our documentation to ensure it remains accurate and usable, which can definitely shift as we see updates to software, such as disaster recovery software we use for replicating servers in the cloud. We see minor updates there all the time. Those minor updates may result in a different set of steps to complete a familiar process. Bottom line, periodic testing ensures plans can still be executed as designed and documented.

Now, testing our plans should include validation of collaborative elements between internal teams and our external vendors. So, the coordination here will be important to ensure that any capabilities or resources we're expecting from those vendors or regulators are available when we have a response condition. The frequency of testing will vary based on the risk profile, the regulations we're under, or the rate of change. If we have a major change in a system or application or process, that would trigger an immediate need for testing. Higher criticality requires more frequent validation. The outcome here is essentially just ensuring each test provides insights for plan improvement and verifies that our plans still work. So, from a scoping perspective, periodic testing ensures the plans remain relevant and effective as conditions evolve. Whether the technology is changing, we have new security threats on the horizon. It ensures our plans are still relevant. Documenting our test outcomes is important because it helps us to refine our incident response plan and recovery strategy. And the scope's going to vary. A targeted component test will be more frequent than a full enterprise-wide drill. So, we'll align tests with the organization's most significant risks. You know, we'll see a full enterprise-wide failover test, for example, maybe once a year, but we'll see smaller read-through type tests even on a monthly basis. So, the frequency principle: less disruptive tests, the read-through, the walkthrough, they're going to happen more often. A lot of environments I see them happening monthly. More disruptive, the full interruption tests are going to be less often, often annually. The improvement loop, we're going to inform management of our results, our weaknesses. That may require revising our plans based on lessons learned. Maybe asking management for additional resources based on what we came back with in those tests.

So, effective testing requires appropriate prioritization and coverage. So, we need to identify critical business applications and supporting infrastructure. We want to prioritize using a risk-based approach, focusing on high-impact systems. And our test infrastructure should cover networks, servers, and applications. We're looking across the service as a whole. Our goal here is to validate backup and failover procedures and ensure recovery within recovery time objectives. So, to truly test our performance against that RTO means at some point we're testing the infrastructure end-to-end, as well as the people and processes and procedures in place to perform that recovery.

So, know these five disaster recovery test types for exam day. There's the read-through test, which is the simplest. You basically distribute the plan and the team reviews it on their own time. Basically checks familiarity. It helps us to find personnel gaps. It allows the owner of each role within the recovery process to review their part of the plan and to update anything they see missing or that has changed. Then there's the walkthrough, what we call a tabletop exercise. The team gathers, they discuss roles and responses for a specific scenario. It is generally moderator-led, and often only the moderator knows the scenario that the team is going to review in the walkthrough in advance. It tests understanding and communication flow, that the team can work together to resolve an incident or recover from a disaster. For example, simulation. Now, this is important. We're stepping from a pure talking situation with read-through and walkthrough to now some doing. So, a simulation is going to be more active. It will test specific components or functions. It may activate some alternate resources, non-critical resources. Generally speaking, a parallel test where we relocate personnel to an alternate site and activate procedures. This can test backup systems running in parallel with production. And so we may have those backup systems running and some folks using those while others remain on standard production systems. And then full interruption, where we shut down the primary site and shift operations fully to the recovery site. This is the most comprehensive test. It validates end-to-end plan effectiveness. It is going to happen less often, often annually. But remember, full interruption is the ultimate proof the plan works.

So, let's talk about the three main categories of recovery testing mentioned by ISACA. There are what they call paper tests. So, the on-paper walkthroughs and discussions done first to refine the plan theoretically. Preparedness tests, localized simulations using actual resources. These test specific parts incrementally. And full operational tests, the near-live tests mimicking disaster. This validates the entire plan execution after lower-level tests have proven successful.

And then evaluating the results. So, recovery tests should produce an expected set of outcomes so that the organization can measure its performance against predefined objectives, like that recovery time objective. So, we need to verify and evaluate plan completeness and accuracy, personnel performance and role execution, the effectiveness of our training and our user awareness. There may be a need for additional training based on personnel performance and role execution. Team and supplier coordination are the internal and external resources collaborating effectively. Are the resources we need there from the external backup site capacity and capability? Is our backup site sufficient? Do we have all the right elements in place to fail over to that recovery site? Vital records retrieval. We want to make sure that any important information we need in a disaster situation or an incident is available to us, which means keeping some critical records offsite and potentially in hard copy. Equipment and supply availability, and then overall operational performance during the test.

So, a few recovery test metrics that would quantify success, identifying needs for improvement, potentially, would include timing and duration, you versus our recovery time objective, accuracy. So, are we recovering with full data integrity? Our configurations intact and secure? Resource utilization. Do we have the staffs and systems in place that we need, or do we need to make some adjustments there? Error issue tracking. So, the number of errors or issues, the severity, the resolution time. Communication effectiveness. Is the speed and clarity of our communication dialed in to our audiences? User and stakeholder satisfaction. So, customers, partners, management, our stakeholders, are they all happy with the response and the outcome? The key takeaway here, analyze metrics systematically and pinpoint weaknesses and refine plans where the metrics tell us we need to.

All right, let's wrap up with a practice question here. So, which of the following provides the best confirmation that the BCDR objectives were achieved during testing? Is it a) better adherence to policies? b) information assets have been assigned to owners per BCDR plans? c) our RPO was exceeded in DR tests? or d) RTO was met during testing? This is a pretty easy question if we just pay attention closely to what they're asking. So, they're asking best confirmation that objectives were met during our BCDR testing. Though we can take better adherence to policies off. That's nice, but it's not a core measurable outcome. Information assets have been assigned to owners per BCDR plan. So, we do that well ahead of any DR test. So, that brings us down to RPO was exceeded in DR test and RTO was met during testing. One clear answer there, and that is recovery time objective was met during testing. So, achieving that RTO provides objective evidence that business continuity disaster recovery objectives have been met. And the indication that the RPO was exceeded means we didn't meet our objective there. For example, if our RPO was 4 hours of data or less, and we went beyond 4 hours of data and maybe we lost 6 hours of data, that means we exceeded, so we didn't meet our RPO. But the RTO in this case is the best indicator of success.

All right, my friends, that does it for domain for part A. I hope you're getting value from the series. As always, if you have questions, drop them in the comments, ping me on LinkedIn. I expect to have domain 4 part B, the conclusion to the series, out sometime in the next 7 days. So, I'll look forward to the big finish with you. And until next time, take care and stay safe. [Music]