Seed-stage B2B SaaS startup
Sample for a fictional organisation · 2,875 words[Company] Disaster Recovery & Business Continuity Plan
- Version: 1.0
- Owner: Founder/CTO
- Approved by: Chief Executive Officer
- Effective date: [Effective date]
- Next review date: [Review date]
1. Purpose
This plan sets out how [Company] keeps its business running and restores its systems and data after a serious disruption, such as a cloud infrastructure outage, a supplier failure, a cyberattack, or an event that prevents employees from working. It is intended to let [Company] continue serving its customers, or resume doing so quickly, with minimal loss of data and minimal disruption to its 8-person team.
This plan sits beneath [Company]'s information security policy. Where a disruption is caused by a security incident, the company's incident response plan governs investigation, evidence handling, and any regulatory or customer notification. This plan governs the continuity of business operations and the recovery of systems and data, and does not repeat notification timelines set out elsewhere.
2. Scope
This plan applies to all systems, data, and suppliers that support [Company]'s product and day-to-day operations, including its infrastructure hosted on AWS, its Google Workspace email and file storage, its GitHub source code repositories, and its remote working environment. It applies to all employees, regardless of location, since the company has no physical office.
Read the full example
[Company] Disaster Recovery & Business Continuity Plan
- Version: 1.0
- Owner: Founder/CTO
- Approved by: Chief Executive Officer
- Effective date: [Effective date]
- Next review date: [Review date]
1. Purpose
This plan sets out how [Company] keeps its business running and restores its systems and data after a serious disruption, such as a cloud infrastructure outage, a supplier failure, a cyberattack, or an event that prevents employees from working. It is intended to let [Company] continue serving its customers, or resume doing so quickly, with minimal loss of data and minimal disruption to its 8-person team.
This plan sits beneath [Company]'s information security policy. Where a disruption is caused by a security incident, the company's incident response plan governs investigation, evidence handling, and any regulatory or customer notification. This plan governs the continuity of business operations and the recovery of systems and data, and does not repeat notification timelines set out elsewhere.
2. Scope
This plan applies to all systems, data, and suppliers that support [Company]'s product and day-to-day operations, including its infrastructure hosted on AWS, its Google Workspace email and file storage, its GitHub source code repositories, and its remote working environment. It applies to all employees, regardless of location, since the company has no physical office.
3. Policy
[Company] must maintain, keep current, and test this plan so that critical systems and data can be recovered within the objectives set out in Section 7, and so that the business can continue operating during a disruption using the arrangements described in this plan. This plan also supports the evidence [Company] needs for its SOC 2 report by documenting its business continuity and disaster recovery practices.
4. Roles and Responsibilities
Given the size of [Company]'s team, this plan uses a small number of roles rather than a formal committee structure. The Founder/CTO leads recovery in most scenarios; a designated deputy stands in when the Founder/CTO is unavailable; and the most senior business decision, such as whether to pay a ransom, is reserved for the Chief Executive Officer.
- Recovery Lead (Founder/CTO): Owns this plan, decides whether to activate it (except where reserved below), leads technical recovery, coordinates with suppliers, and communicates with customers and employees during a disruption.
- Deputy Recovery Lead ([Deputy Recovery Lead name/role]): Performs the Recovery Lead's duties if the Recovery Lead is unreachable or unable to act.
- Chief Executive Officer: Decides whether to pay a ransom demand, approves major customer communications during a significant disruption, and acts as Recovery Lead if both the Recovery Lead and Deputy Recovery Lead are unavailable.
- All employees: Follow instructions issued during an activation, use the fallback communication channel in Section 6 if needed, and report any disruption to the Recovery Lead as soon as they notice it.
5. Plan Activation
This plan may be activated by the Recovery Lead. If the Recovery Lead is unreachable, the Deputy Recovery Lead or the Chief Executive Officer may activate it instead. Activation should be based on the following criteria:
- A disruption to any system or service is expected to last longer than the recovery time objective set for it in Section 7.
- Loss of access to a supplier, tool, or account needed to deliver [Company]'s product or serve customers, for longer than 24 hours.
- A confirmed security incident, including ransomware, that affects the availability of production systems or customer data.
- An event that prevents more than half of [Company]'s employees from working for more than one calendar day.
6. Communications and Escalation
The primary channels for internal communication during a disruption are the team chat tool and email. Because email may run on Google Workspace, which could itself be affected by an outage, the fallback channel is a phone call or text message to employees' mobile phones, since this does not depend on any of [Company]'s cloud accounts or suppliers. The Recovery Lead must maintain an up-to-date list of employee mobile numbers outside the systems this plan protects, as described in Appendix A.
Customers must be told promptly about any disruption that affects the customer-facing application or their data, using email or another channel the customer has agreed to. The Recovery Lead must check the affected customer's contract for any availability commitment or notice period before sending or finalizing the communication, since these vary by customer.
7. Recovery Time and Recovery Point Objectives
A recovery time objective (RTO) is the longest a system or service may be unavailable, measured from the start of the disruption, before the company must have it working again. A recovery point objective (RPO) is the most data, measured in time, that [Company] can afford to lose, meaning the maximum acceptable gap between the last usable backup and the point of failure. Where a system is run by a supplier that [Company] cannot restore itself, the RTO is the longest [Company] can work without that system before switching to a workaround.
| System or service | Recovery time objective | Recovery point objective |
|---|---|---|
| Customer-facing application and database (AWS) | 24 hours | 1 hour |
| GitHub source code repositories | 24 hours | 24 hours |
| Google Workspace email and file storage | 24 hours | 24 hours |
Because [Company] has 8 employees and no on-call rota, any of these RTOs may fall outside normal business hours or across a weekend. The Recovery Lead must be reachable by phone call to their mobile number at any time to meet these objectives; the Deputy Recovery Lead must be reachable in the same way if the Recovery Lead cannot be.
8. Backup and Restoration
[Company] must maintain backups that are made at least as often as the recovery point objectives in Section 7, with at least one copy of each backup kept separate from production systems and accounts, so that an attacker or an administrator error affecting production cannot also delete or encrypt the backup.
- Automated backups of the production database must be taken at least once every hour and retained for at least 30 days.
- The production application's infrastructure configuration must be stored in version control and mirrored to a storage location outside AWS at least once every 24 hours, retained for at least 90 days.
- GitHub repositories must be cloned or mirrored to a storage location outside GitHub at least once every 24 hours and retained for at least 90 days.
- Google Workspace email and file storage must be exported or backed up to a storage location outside Google Workspace at least once every 24 hours and retained for at least 30 days.
- Restoration of each critical system listed in Section 7 must be tested at least once per year, with results recorded as set out in Section 12.
- Backups containing personally identifiable information must be protected with the same access controls as the production data they are copied from.
9. Continuity of Critical Services
[Company]'s core service runs on AWS, and the company is responsible for restoring that infrastructure and its data itself, following the backup and restoration requirements in Section 8. If AWS experiences an outage affecting the region [Company] uses, the company must be able to restore its infrastructure into a different AWS region from its backups, since it operates no standing infrastructure elsewhere.
Google Workspace and GitHub are run by their respective providers, and [Company] depends on those providers to restore their own services. Those providers may not be able to recover data that [Company]'s own users deleted, or that an attacker encrypted or deleted, so [Company] must hold its own separate copy of the data it cannot afford to lose from each of these tools, as set out in Section 8, rather than relying solely on the provider.
10. Alternate Work Facilities
[Company] has no physical office; all employees work remotely. If an employee loses power or internet access at home, they must use a mobile phone, a mobile hotspot, or an alternate location such as a coworking space or public internet access point to continue working, and must notify the Recovery Lead using the fallback channel in Section 6.
If an event affects multiple employees at once, such as a regional power or network outage, the Recovery Lead must assess how many employees are affected and adjust expectations for response and recovery accordingly, communicating any resulting delay to customers as described in Section 6.
11. Business Continuity Procedures by Scenario
11.1 Customer-Facing Application Unavailable
This scenario covers unavailability of [Company]'s customer-facing application hosted on AWS due to an infrastructure failure, application error, or capacity issue, where the cause is not a security incident.
- Monitoring or a customer report alerts the Recovery Lead that the application is unavailable.
- The Recovery Lead assesses the scope, cause, and expected duration of the outage using AWS status information and internal monitoring.
- If the outage is expected to exceed the recovery time objective in Section 7, the Recovery Lead activates this plan under Section 5.
- The Recovery Lead restores the affected service from the most recent backup or infrastructure configuration, following Section 8.
- The Recovery Lead verifies that the application is functioning correctly and that customer data is intact before declaring the incident resolved.
- The Recovery Lead notifies affected customers of the outage and its resolution, per Section 6.
- The Recovery Lead records the event for review under Section 12.
11.2 Internal Applications Inaccessible
This scenario covers loss of access to internal tools [Company] depends on for its own operations, such as Google Workspace email and file storage or the GitHub source code repository, where the disruption originates with the supplier rather than [Company]'s own infrastructure.
- The employee who first notices the disruption reports it to the Recovery Lead, using the fallback channel in Section 6 if the primary channel is affected.
- The Recovery Lead checks the supplier's status page and support channels to confirm the outage and estimate its duration.
- If the expected outage exceeds the recovery time objective for the affected tool in Section 7, the Recovery Lead activates this plan and directs employees to the fallback channel and to the separate backup copies described in Section 8.
- Employees continue essential work using the fallback channel and the available backup copies until the supplier restores service.
- The Recovery Lead confirms restoration of service and checks that no data was lost beyond the recovery point objective in Section 7.
- The Recovery Lead records the event for review under Section 12.
11.3 Cybersecurity Breach Recovery
This scenario covers restoring business operations after a confirmed security incident. Investigation, evidence handling, and any regulatory or customer notification about the breach are governed by [Company]'s incident response plan; this plan governs only the continuity and recovery of affected systems and services.
- The Recovery Lead confirms with whoever is leading the incident response that affected systems have been contained and are safe to recover.
- The Recovery Lead assesses which business services are disrupted and whether the disruption is expected to exceed the recovery time objective in Section 7.
- If so, the Recovery Lead activates this plan under Section 5.
- The Recovery Lead restores affected systems from a backup confirmed to predate the incident, following Section 8.
- The Recovery Lead verifies system integrity and functionality before returning the system to production use.
- The Recovery Lead communicates with affected customers as needed, per Section 6.
- The Recovery Lead records the event for review under Section 12.
11.4 Ransomware Attack
This scenario covers recovery from ransomware or similar malware that encrypts or destroys company data. This plan governs isolation and restoration; the incident response plan governs investigation and notification.
- The person who discovers the ransomware immediately disconnects the affected system(s) from the network to stop it spreading, and reports it to the Recovery Lead using the fallback channel in Section 6 if needed.
- The Recovery Lead isolates any other systems that may be affected and notifies whoever is leading the incident response process.
- The Recovery Lead identifies the most recent backup confirmed to be clean and unaffected by the ransomware before restoring anything.
- The Recovery Lead restores affected systems only from that clean backup, following Section 8.
- The Chief Executive Officer decides whether to pay any ransom demand, after obtaining legal advice and confirming that payment would not breach sanctions law; this decision is not delegated to any other role.
- If [Company] holds cyber insurance, the Recovery Lead or Chief Executive Officer notifies the insurer promptly.
- The Recovery Lead records the event for review under Section 12.
11.5 Natural Disaster or Loss of a Site or Region
This scenario covers loss of the AWS region hosting [Company]'s infrastructure, or an event that prevents employees from working from their home location, such as a regional power outage or natural disaster.
- The Recovery Lead determines whether the disruption affects the AWS infrastructure, individual employees, or both.
- If the AWS region is affected, the Recovery Lead works with AWS support to assess restoration timelines and, if the outage is expected to exceed the recovery time objective in Section 7, restores infrastructure into a different AWS region from backups, following Section 8.
- If employees are affected, they use the fallback communication channel in Section 6 and, where possible, an alternate location or connection to continue working, per Section 10.
- If the Recovery Lead is personally affected and unavailable, the Deputy Recovery Lead assumes the role.
- The Recovery Lead confirms restoration of service and communicates with customers if the disruption affected the customer-facing application, per Section 6.
- The Recovery Lead records the event for review under Section 12.
11.6 Failure of a Critical Supplier
This scenario covers unavailability of a supplier [Company] depends on to deliver its product or run its business, including AWS for infrastructure, Google Workspace for email and file storage, and GitHub for source code hosting.
- The Recovery Lead identifies the affected supplier and checks its status page and support channels for an estimated restoration time.
- If the estimated outage is within the recovery time objective for the affected service in Section 7, the Recovery Lead monitors the situation and keeps employees updated.
- If the outage is expected to exceed the recovery time objective, the Recovery Lead activates this plan and directs employees to the applicable workaround or backup copy described in Section 8.
- If the outage is extended or repeated, the Recovery Lead evaluates whether an alternate supplier is needed and raises this with the Chief Executive Officer.
- The Recovery Lead communicates any customer impact per Section 6.
- The Recovery Lead records the event for review under Section 12.
12. Testing and Review
[Company] must run a tabletop exercise of this plan at least once per year, in which the Recovery Lead and Deputy Recovery Lead walk through at least one scenario from Section 11. Restore testing under Section 8 must be carried out at least once per year for each critical system. This plan must be reviewed and updated after every activation, after any significant change to [Company]'s systems or suppliers, and at least once per year regardless of whether either of those events has occurred.
13. Appendix A: Emergency Contact and Recovery Information Template
This plan and the contact list below must remain reachable even when [Company]'s main systems are down, for example as a printed copy or a copy stored outside Google Workspace, AWS, and GitHub.
| Role or party | Name | Contact details | Notes |
|---|---|---|---|
| Recovery Lead (Founder/CTO) | [Recovery Lead name] | [Recovery Lead phone/email] | Primary activator of this plan |
| Deputy Recovery Lead | [Deputy Recovery Lead name] | [Deputy Recovery Lead phone/email] | Stands in if Recovery Lead is unavailable |
| Chief Executive Officer | [CEO name] | [CEO phone/email] | Decides ransom payment; escalation point |
| AWS support | N/A | [AWS support contact / account ID] | Infrastructure hosting provider |
| Google Workspace support | N/A | [Google Workspace support contact] | Email and file storage provider |
| GitHub support | N/A | [GitHub support contact] | Source code hosting provider |
| Cyber insurer | [Insurer name, if applicable] | [Insurer contact details] | Notify promptly after a ransomware or breach event |
| External legal counsel | [Law firm or attorney name] | [Legal counsel contact details] | Advises on ransom payment and sanctions law |
| Key customer contacts | N/A | [Reference to customer contact list location] | Used for customer notifications under Section 6 |
Disclaimer
This document is provided for informational purposes only and does not constitute legal advice. It is provided "as is", without warranty of any kind, express or implied, and no liability is accepted for any loss or damage arising from its use. It is used at your own discretion. Review it with a qualified adviser before adopting it.