Going multi-region is one of those decisions that sounds straightforward on a whiteboard and becomes deeply technical the moment you start. "Run the application in two regions" turns into a year of work on data replication, DNS, certificate management, and operations runbooks. The patterns described here are the ones production systems actually use — not the textbook diagrams.
There are three reasons to go multi-region, and they are not the same. Knowing which one applies decides almost every downstream choice.
The Three Reasons
1. Disaster recovery. You need a second region you can fail to if the primary becomes unreachable. Traffic normally goes to one region; the other exists to take over.
2. Latency. You serve users globally and a single region cannot deliver the experience you need. Each user is routed to the nearest region.
3. Compliance. Data for users in a given jurisdiction must be stored and processed in that jurisdiction. Each region is the source of truth for its users.
The data architecture, the DNS routing, the failover process, and the cost profile differ between these three. Trying to design a multi-region system without picking one is what produces architecture astronauts.
The Four Architectural Patterns
Active-Passive
One region serves traffic; the other replicates data and waits. Failover is a deliberate event, often manual.
Region A (active) ────► All user traffic
│
│ data replication
▼
Region B (passive) ─── (no traffic)
Best for disaster recovery. Cheap in steady state. Failover is slow (minutes to an hour) and disruptive.
Active-Active with Region Affinity
Both regions serve traffic, but each user is sticky to one region. Cross-region writes are rare.
Users (US) ─────────► Region A (US data)
↕ small cross-region replication
Users (EU) ─────────► Region B (EU data)
Best for latency or compliance. Failure of one region affects only its users. Cross-region replication exists for disaster recovery, not for serving traffic.
Active-Active with Global Writes
Both regions accept writes for the same data. Conflicts are resolved through coordination or eventual consistency.
Users ─────────► Region A ⇄ Region B ⇐───── Users
↕ conflict resolution
The most powerful, the most expensive. You need a strategy for conflicting writes — strong consistency through Spanner-style global coordination, last-write-wins, CRDTs, or application-specific merge logic.
Hub and Spoke
One region is the system of record; others are read-only or partial replicas.
┌──── Edge region (read-only) ────► users
│
Source region ──┼──── Edge region (read-only) ────► users
│
└──── Edge region (read-only) ────► users
Best for read-heavy global workloads. Writes pay the cost of going to the source region; reads are local. CDNs and edge caches generalize this pattern.
The Database Decision Dominates Everything
The pattern you can adopt is determined by your database. Most multi-region pain is database pain.
Relational with leader-follower replication (PostgreSQL, MySQL): natural fit for active-passive. Writes go to the leader. Followers can serve reads but have replication lag. Promoting a follower to leader is a manual or scripted event.
Spanner / Yugabyte / CockroachDB: designed for active-active global writes. Strong consistency across regions, at the cost of higher write latency (cross-region consensus).
DynamoDB Global Tables / Cassandra: active-active with last-write-wins or application-level conflict resolution. High write throughput, eventual consistency.
Append-only logs (Kafka, etc.): can be mirrored across regions with MirrorMaker or similar. Conflicts are avoided by partitioning — each region owns a subset of partitions.
Object storage (S3) cross-region replication: asynchronous, simple, and built into the storage layer. Good fit when objects are immutable.
Pick the database based on the pattern you need, not the pattern based on the database you have. Migrating off the wrong database is a multi-year project.
DNS and Traffic Routing
Once data is replicated, traffic has to find the right region.
Latency-based routing. DNS returns the region nearest to the resolver. Route 53, Cloud DNS, NS1 all support it. Works for active-active with region affinity. Failure detection is slow (DNS TTLs and health checks).
Geo routing. DNS returns based on the resolver's country or region, not latency. Best for compliance — EU users to EU regions even if the network is faster to a US region.
Anycast. A single IP is announced from multiple regions; BGP routes the user to the nearest one. Powerful but only available from certain providers (Cloudflare, Fastly, AWS Global Accelerator).
Application-level routing. The frontend decides which API region to call based on the user's account, regardless of where they are physically. Essential for compliance scenarios.
Most production multi-region systems use a combination: anycast or latency-based for the public edge, then application-level routing for authenticated traffic to the correct user-data region.
Health Checks and Failover
The textbook says "if region A is unhealthy, route to region B." The hard part is deciding what unhealthy means.
- Synthetic health checks. A probe hits an endpoint every few seconds. If failures cross a threshold, DNS or the load balancer fails over. Works for total outages.
- Application-level health. The endpoint runs lightweight queries against the database, cache, and downstream services. Detects more failure modes but has more false positives.
- Dependency health. A region is considered healthy only if its database, cache, and key downstream services are healthy. The most accurate, the most complex.
Beware of flapping: a region that oscillates between healthy and unhealthy causes more damage than one that is clearly down. Hysteresis (require N consecutive failures to fail over, and M consecutive successes to fail back) is essential.
The Operations Reality
Multi-region adds operational concerns that are easy to miss:
- Certificates. Each region's load balancer needs valid TLS certificates. ACME or your CA has to be set up per region.
- Secrets. Secrets must be accessible from every region or replicated to each.
- Observability. Logs, metrics, and traces from every region have to land in a place where you can correlate across them.
- Deployments. Releases have to happen in every region. Blue/green and canary deployments multiply in complexity.
- Cost. Cross-region data transfer is the single largest cost surprise. Egress between AWS regions runs $0.02/GB. Across millions of requests, this adds up.
Honest Advice
Most teams should not be multi-region. A single, well-operated region with good backups handles 99.9% availability cheaper than two regions with bad operations. Going multi-region for "high availability" rarely pays back unless the team is mature enough to run the additional complexity well.
Go multi-region when:
- A regulator requires data residency
- A real revenue impact comes from latency to a specific geographic market
- You have a measurable need for a region-loss disaster recovery option, and your team is ready to operate it
Otherwise, invest in single-region reliability first. It's cheaper and you get most of the benefit.
Weighing whether your next reliability investment should be multi-region or single-region done right? We help teams scope reliability work to the actual revenue at risk. scopeforged.com