Skip to content
Web Development

Ecommerce Disaster Recovery: How to Plan for Outages, Data Loss and Rollback

How to plan ecommerce disaster recovery: failure scenarios, RPO and RTO, backups, database recovery, infrastructure, integrations, DNS, payments, rollback and DR testing.

Quick answer

Ecommerce disaster recovery starts with two numbers per system: the recovery point objective (how much data you can lose, measured in time) and the recovery time objective (how long you can be down). Choose a strategy to meet them (backup and restore, pilot light, warm standby or active-active), back up databases, media, configuration and platform data, plan for integrations, DNS and payments, keep fast rollback for bad deployments, and test restores and failovers regularly. SaaS platforms reduce, but do not remove, your DR responsibilities.

Where This Fits

Detection is covered in observability and scaling in scalability. Launch rollback planning is covered in migration launch plan, and security incidents in ecommerce security.

Failure Scenarios

ScenarioExampleTypical response
Bad deploymentRelease breaks checkoutRoll back; fix forward later
Data corruption or deletionScript overwrites pricesPoint-in-time restore of affected data
Infrastructure outageHosting region unavailableFail over to another region or restore
Third-party failurePayment, search or tax provider downDegrade gracefully, use fallback
Security incidentCompromised credentials, ransomwareIsolate, restore from clean backups
DNS or certificate problemDomain misconfigurationPre-documented DNS rollback
Integration failureERP unavailable for daysQueue and reconcile

RPO and RTO Explained

RPO is about data: if your last usable backup was an hour ago, an outage now loses up to an hour of orders and changes, so your RPO is one hour. RTO is about time: if restoring takes four hours, your RTO is at least four hours. Set both per system from business impact. Order and payment data usually needs a very low RPO; a content CMS can tolerate more.

SystemExample RPOExample RTO
Orders and paymentsMinutes or lessUnder an hour
Catalog and pricingHoursA few hours
Customer accountsMinutes to hoursA few hours
Content and mediaA dayA day
AnalyticsA day or moreDays

Worth noting

These values are illustrative. Set your own from the cost of downtime and lost data for your business.

Recovery Strategies

Cloud provider guidance, such as the AWS Well-Architected disaster recovery whitepaper, describes four broad strategies:

StrategyHow it worksRecovery profileCost
Backup and restoreRestore from backups into new infrastructureSlowest; RPO depends on backup frequencyLowest
Pilot lightCore data replicated, minimal infrastructure ready to scaleFasterLow
Warm standbyScaled-down full copy running elsewhereFaster stillMedium
Active-active (multi-site)Serving from several locations at onceFastest, near-zero data loss possibleHighest

Not sure how long your store would be down after a major failure?

ZSpace can set RPO and RTO with you, map current gaps and design a recovery plan proportionate to what downtime costs you.

Start a Project

Backups

  • Automated database backups with point-in-time recovery where available
  • Media and file storage backed up or versioned
  • Infrastructure as code and configuration in version control
  • Secrets backed up securely
  • Backups stored in a separate account or region
  • Immutable or protected backups against ransomware
  • SaaS platform data exported regularly (products, customers, orders, themes)

Database Recovery

Know how to restore a full database and how to recover specific data, such as prices overwritten by a bad import, without rolling back everything else. Practise both. Restoring an entire database to fix one table can lose hours of orders.

Infrastructure

Define infrastructure as code so environments can be recreated. Use redundancy within a region for common failures, and decide whether cross-region recovery is justified by your RTO. Keep capacity limits and quotas in the recovery region sufficient.

Integrations

Integrations complicate recovery: restoring a database to an earlier point may cause it to disagree with the ERP, payment provider or warehouse. Plan reconciliation after recovery, keep queues durable so messages are not lost during outages, and know which systems are sources of truth for orders and payments. See queue architecture.

DNS

Document DNS records, keep TTLs appropriate for planned failovers, secure registrar and DNS provider accounts with strong authentication, and practise switching traffic. DNS mistakes during recovery can extend an outage.

Payment Considerations

During outages, payments may be authorized without orders being created, or orders created without confirmed payment. Reconcile payment provider records with orders after recovery, contact affected customers, and avoid double charges. Consider whether a backup payment method or provider is worth configuring for critical periods.

Rollback

Bad deployments are frequent incidents. Keep the previous release ready to redeploy, use feature flags to switch off risky features, make database migrations backward compatible so rollback does not require data changes, and avoid deploying just before peak events.

Testing

  • Regular restore tests from backups, timed against RTO
  • Point-in-time recovery tests for critical tables
  • Failover exercises where you have standby infrastructure
  • Rollback drills for deployments
  • Tabletop exercises for security and provider failures
  • Documented results and fixed gaps

Worked Example

An illustrative scenario, not a client case: a bulk price import script overwrites thousands of prices with incorrect values during trading hours. The store has nightly backups but no point-in-time recovery, so restoring the whole database would lose a day of orders. After fixing the immediate problem by re-importing from the source file, the team enables point-in-time recovery, adds a review step for bulk imports, practises restoring a single table into a separate database and documents the runbook.

Implementation Steps

  • Agree RPO and RTO per system
  • Choose a recovery strategy per system
  • Implement protected, separate backups
  • Write runbooks for the top failure scenarios
  • Plan integration and payment reconciliation
  • Test restores, failover and rollback on a schedule

Common Mistakes

  • Backups that have never been restored
  • Backups in the same account as production
  • No RPO and RTO agreed
  • Ignoring integrations and payments in recovery plans
  • Assuming the SaaS platform handles everything
  • Runbooks nobody has used

Conclusion

Disaster recovery works when RPO and RTO are agreed per system, the strategy matches them, backups are protected and restorable, integrations and payments are reconciled, rollback is fast and everything is tested. Related: observability and scalability.

FAQ

Common questions

The plans, systems and practices that let a store recover from serious failures (infrastructure outages, data loss or corruption, security incidents, bad deployments, provider failures) within agreed time and data-loss limits.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.