Ecommerce Disaster Recovery: How to Plan for Outages, Data Loss and Rollback
How to plan ecommerce disaster recovery: failure scenarios, RPO and RTO, backups, database recovery, infrastructure, integrations, DNS, payments, rollback and DR testing.
Quick answer
Ecommerce disaster recovery starts with two numbers per system: the recovery point objective (how much data you can lose, measured in time) and the recovery time objective (how long you can be down). Choose a strategy to meet them (backup and restore, pilot light, warm standby or active-active), back up databases, media, configuration and platform data, plan for integrations, DNS and payments, keep fast rollback for bad deployments, and test restores and failovers regularly. SaaS platforms reduce, but do not remove, your DR responsibilities.
Where This Fits
Detection is covered in observability and scaling in scalability. Launch rollback planning is covered in migration launch plan, and security incidents in ecommerce security.
Failure Scenarios
| Scenario | Example | Typical response |
|---|---|---|
| Bad deployment | Release breaks checkout | Roll back; fix forward later |
| Data corruption or deletion | Script overwrites prices | Point-in-time restore of affected data |
| Infrastructure outage | Hosting region unavailable | Fail over to another region or restore |
| Third-party failure | Payment, search or tax provider down | Degrade gracefully, use fallback |
| Security incident | Compromised credentials, ransomware | Isolate, restore from clean backups |
| DNS or certificate problem | Domain misconfiguration | Pre-documented DNS rollback |
| Integration failure | ERP unavailable for days | Queue and reconcile |
RPO and RTO Explained
RPO is about data: if your last usable backup was an hour ago, an outage now loses up to an hour of orders and changes, so your RPO is one hour. RTO is about time: if restoring takes four hours, your RTO is at least four hours. Set both per system from business impact. Order and payment data usually needs a very low RPO; a content CMS can tolerate more.
| System | Example RPO | Example RTO |
|---|---|---|
| Orders and payments | Minutes or less | Under an hour |
| Catalog and pricing | Hours | A few hours |
| Customer accounts | Minutes to hours | A few hours |
| Content and media | A day | A day |
| Analytics | A day or more | Days |
Worth noting
These values are illustrative. Set your own from the cost of downtime and lost data for your business.
Recovery Strategies
Cloud provider guidance, such as the AWS Well-Architected disaster recovery whitepaper, describes four broad strategies:
| Strategy | How it works | Recovery profile | Cost |
|---|---|---|---|
| Backup and restore | Restore from backups into new infrastructure | Slowest; RPO depends on backup frequency | Lowest |
| Pilot light | Core data replicated, minimal infrastructure ready to scale | Faster | Low |
| Warm standby | Scaled-down full copy running elsewhere | Faster still | Medium |
| Active-active (multi-site) | Serving from several locations at once | Fastest, near-zero data loss possible | Highest |
Not sure how long your store would be down after a major failure?
ZSpace can set RPO and RTO with you, map current gaps and design a recovery plan proportionate to what downtime costs you.
Backups
- Automated database backups with point-in-time recovery where available
- Media and file storage backed up or versioned
- Infrastructure as code and configuration in version control
- Secrets backed up securely
- Backups stored in a separate account or region
- Immutable or protected backups against ransomware
- SaaS platform data exported regularly (products, customers, orders, themes)
Database Recovery
Know how to restore a full database and how to recover specific data, such as prices overwritten by a bad import, without rolling back everything else. Practise both. Restoring an entire database to fix one table can lose hours of orders.
Infrastructure
Define infrastructure as code so environments can be recreated. Use redundancy within a region for common failures, and decide whether cross-region recovery is justified by your RTO. Keep capacity limits and quotas in the recovery region sufficient.
Integrations
Integrations complicate recovery: restoring a database to an earlier point may cause it to disagree with the ERP, payment provider or warehouse. Plan reconciliation after recovery, keep queues durable so messages are not lost during outages, and know which systems are sources of truth for orders and payments. See queue architecture.
DNS
Document DNS records, keep TTLs appropriate for planned failovers, secure registrar and DNS provider accounts with strong authentication, and practise switching traffic. DNS mistakes during recovery can extend an outage.
Payment Considerations
During outages, payments may be authorized without orders being created, or orders created without confirmed payment. Reconcile payment provider records with orders after recovery, contact affected customers, and avoid double charges. Consider whether a backup payment method or provider is worth configuring for critical periods.
Rollback
Bad deployments are frequent incidents. Keep the previous release ready to redeploy, use feature flags to switch off risky features, make database migrations backward compatible so rollback does not require data changes, and avoid deploying just before peak events.
Testing
- Regular restore tests from backups, timed against RTO
- Point-in-time recovery tests for critical tables
- Failover exercises where you have standby infrastructure
- Rollback drills for deployments
- Tabletop exercises for security and provider failures
- Documented results and fixed gaps
Worked Example
An illustrative scenario, not a client case: a bulk price import script overwrites thousands of prices with incorrect values during trading hours. The store has nightly backups but no point-in-time recovery, so restoring the whole database would lose a day of orders. After fixing the immediate problem by re-importing from the source file, the team enables point-in-time recovery, adds a review step for bulk imports, practises restoring a single table into a separate database and documents the runbook.
Implementation Steps
- Agree RPO and RTO per system
- Choose a recovery strategy per system
- Implement protected, separate backups
- Write runbooks for the top failure scenarios
- Plan integration and payment reconciliation
- Test restores, failover and rollback on a schedule
Common Mistakes
- Backups that have never been restored
- Backups in the same account as production
- No RPO and RTO agreed
- Ignoring integrations and payments in recovery plans
- Assuming the SaaS platform handles everything
- Runbooks nobody has used
Ready to make recovery a tested capability?
Talk to ZSpace about resilience and recovery engineering, backup and reconciliation automation and Shopify data protection.
Conclusion
Disaster recovery works when RPO and RTO are agreed per system, the strategy matches them, backups are protected and restorable, integrations and payments are reconciled, rollback is fast and everything is tested. Related: observability and scalability.
Common questions
The plans, systems and practices that let a store recover from serious failures (infrastructure outages, data loss or corruption, security incidents, bad deployments, provider failures) within agreed time and data-loss limits.