← All servicesService

High availability on AWS

I design availability around what an outage costs your business: Multi-AZ, automatic failover, and an RTO/RPO decided up front, not after the first incident.

Get an availability assessmentThe quote comes out of the assessment.

01When it makes sense

  • An hour down costs the business more than a month of infrastructure.
  • If one machine goes down, the service does.
  • Nobody knows how long it takes to come back, or how much data is lost.

02How the work goes

  1. 01

    Assess

    What an outage costs; the target RTO/RPO and the single points of failure follow from it.

  2. 02

    Design

    Multi-AZ, load balancing and database failover, written in Terraform.

  3. 03

    Failure test

    An instance and the database are made to fail, and recovery is measured.

  4. 04

    Handoff

    Incident runbooks and alarms in place.

03What is included

  • Amazon RDS Multi-AZ.
  • Application Load Balancer and Auto Scaling across zones.
  • Backups and restores tested against the RPO.
  • Alarms in Amazon CloudWatch.

What is not

  • 24/7 on-call after handoff.

04Records

A 3,000-user peak that held, and a platform that replaces a failed node on its own.

05Questions

Do I need multi-region?
Only if a regional outage costs more than running it. For most, Multi-AZ is enough, and the assessment says so with numbers.
How much does it cost?
It comes out of the assessment: with the inventory I put together a quote, and you approve it before anything starts. What AWS bills is pay-per-use and goes straight to AWS.
How do I know it works?
Because it is tested: failure is forced before handoff and recovery is measured.

Tell me what you have today and I will reply, almost always within 24 hours.

Get an availability assessment