Business Continuity and Disaster Recovery Toolkit

Collection(s): EDUCAUSE Working Group

Planning

The best plan is a simple one. Starting your planning from scratch or working from a significantly outdated plan can be intimidating. Try not to overthink it! If you have specialized software to help lay out your plan, great! If not, start simple, as described below. Start small, set up regularly occurring meetings, and continue to build your plan. Document everything! When developing your BC/DR plan, the questions you need to start with are the simplest: who, what, and how. Beyond that, you will need to establish recovery time objectives (RTOs) and recovery point objectives (RPOs).

Use this Disaster Recovery Plan template to guide your planning.

Who

In an emergency, knowing your stakeholders is key. Who has the authority to declare an emergency? What teams need to be included to address the situation? Who needs to be informed regularly? Once stakeholders have been determined, their contact information and areas of responsibility should be documented and readily available.

What

An important question to have consensus on is "what constitutes an emergency?" In the chaos of any incident, it's important to know what conditions need to be met to trigger the BC/DR plan. These conditions may differ depending on the type of disaster, and each should be carefully outlined.

How

Most of your BC/DR plan will answer the "how" questions. How will teams be deployed to address the emergency? How do you maintain systems during an emergency? How do you know an emergency is over? How will you improve your plan for the future?

Develop RTO and RPO Standards

Determining RTOs and RPOs in your BC/DR plan helps establish the urgency of the response and determine which members of the team need to be assembled.

Recovery Time Objective: Duration of time an application can be down without considerably damaging business operations (system focused).

Convene key stakeholders to determine your recovery time objective: service owners, data custodians, functional users, help desk, enterprise risk management, emergency management services.

  1. Brainstorm and list critical enterprise systems:
    1. Which user-facing applications need to be constantly available for our n community?
    2. Which applications are most critical for our mission-focused operations to function?
    3. Which applications are in the front line of security protection?
  2. Document RTO:
    1. Which of these systems can we do without for 72 hours (or longer)?
    2. Which of these systems can we do without for no more than 48 hours?
    3. Which of these systems can we do without for no more than 24 hours?
  3. Prioritize according to RTO for each critical system: If some systems must be brought back online before others, include the appropriate priority order.
  4. Document key mitigations: What tools and processes do we currently have in place that we can use to mitigate the loss of functionality due to system downtime?

Recovery Point Objective: The amount of data that can be lost without considerably damaging business operations or finances (data focused; measures loss tolerance).

  1. Run tests to determine how quickly data needs to be available for each enterprise application based on RTO prioritization.
  2. Categorize critical enterprise systems depending on their backup restoration requirements:
    1. How soon does data need to be restored?
  3. Calculate your financial position regarding backups:
    1. For how many applications can we afford to maintain fail-over and replication services?
    2. Do we need to store some backups offline, on hard drives, or tape?
  4. Prioritize according to data restoration needs (#2) and financial position for recovery.

Use this RTO/RPO Plan template to document your RTO and RPO.