High availability on AWS
I design availability around what an outage costs your business: Multi-AZ, automatic failover, and an RTO/RPO decided up front, not after the first incident.
Get an availability assessmentThe quote comes out of the assessment.
01When it makes sense
- An hour down costs the business more than a month of infrastructure.
- If one machine goes down, the service does.
- Nobody knows how long it takes to come back, or how much data is lost.
02How the work goes
- 01
Assess
What an outage costs; the target RTO/RPO and the single points of failure follow from it.
- 02
Design
Multi-AZ, load balancing and database failover, written in Terraform.
- 03
Failure test
An instance and the database are made to fail, and recovery is measured.
- 04
Handoff
Incident runbooks and alarms in place.
03What is included
- Amazon RDS Multi-AZ.
- Application Load Balancer and Auto Scaling across zones.
- Backups and restores tested against the RPO.
- Alarms in Amazon CloudWatch.
What is not
- 24/7 on-call after handoff.
04Records
A 3,000-user peak that held, and a platform that replaces a failed node on its own.
05Questions
- Do I need multi-region?
- Only if a regional outage costs more than running it. For most, Multi-AZ is enough, and the assessment says so with numbers.
- How much does it cost?
- It comes out of the assessment: with the inventory I put together a quote, and you approve it before anything starts. What AWS bills is pay-per-use and goes straight to AWS.
- How do I know it works?
- Because it is tested: failure is forced before handoff and recovery is measured.
Tell me what you have today and I will reply, almost always within 24 hours.
Get an availability assessment