Core Concept

Replication, Failover, & Leader Election

Replication adds read capacity and failover, but sync vs async replication and leader election determine how much data you might lose or how stale reads can get.


1. What It Is

When you need the database to survive a node crash, replication and failover are the first tools you reach for. One leader accepts writes; replicas catch up; election picks a new leader when heartbeats stop.

What:

Replication clones database states to backup servers. Failover promotes a replica when the primary fails, coordinated by leader election protocols (Raft/Paxos).

Primary purpose:

Ensure high availability, survive hardware crashes, and scale read query throughput globally.

Usually used for:

Primary-replica database clusters, stateful messaging pools, and decentralized consensus networks.

2. Core Mental Model

We picture one leader accepting writes while replicas catch up β€” sync or async β€” and failover promoting a replica when the leader disappears:

🚦 Replication Latency Bounds

Balance consistency: pick Synchronous for financials; choose Asynchronous for read-heavy social profiles.

πŸ—³οΈ Majority Quorum (N/2 + 1)

Leader election require consensus agreement from over 50% of the cluster nodes to prevent partitioned updates.

πŸ”„ Heartbeat Monitors

Deploy strict periodic heartbeat signals to monitor master health. Trigger failover sweeps when nodes drop heartbeats.

In the room

Candidates often say "read replicas" without naming sync vs async. State your RPO: synchronous replication means RPO β‰ˆ 0 but higher write latency; async is faster but readers may see stale data. If they ask about split-brain, mention quorum (N/2 + 1) and fencing.

3. Why It Matters in HLD

Replication is how we survive node death without losing the product. We frame the design around three lenses:

Needed When:

Designing active-passive database tiers, highly-available queues, or resilient coordination rings.

Avoids:

Single-point-of-failure downtimes, permanent data losses from disk crashes, and split-brain dual leaders conflicts.

Optimizes For:

Continuous uptime SLAs, disaster recovery speeds, read throughput capacity, and state consensus stability.

4. Architecture & Data Flow

Narrate the replication path as interview steps. Step 1 β€” Write path: client sends mutation to the primary leader. Step 2 β€” Replicate: leader streams WAL entries to followers β€” sync waits for quorum, async acks immediately. Step 3 β€” Read path: route analytics to replicas; state whether stale reads are acceptable. Step 4 β€” Heartbeat loss: followers miss leader heartbeats and start election. Step 5 β€” Failover: new leader promoted; proxy or DNS cuts traffic to the survivor.

Loading...

5. Key Characteristics

Sync vs async replication and Raft vs Bully are the two tables we walk through on the whiteboard:

  • Sync replication blocks until replicas persist; async acknowledges immediately β€” your RPO/RTO follows from this choice:
Type ModelBehavioral MechanicDurability BenefitLatency Cost
SynchronousPrimary blocks write acknowledgments until secondaries write to disk.Zero data loss window during primary crashes.
  • Write latencies increase
  • if a secondary node fails, all writes block.
Asynchronous
  • Primary acknowledges write instantly
  • syncs secondaries in background.
  • Ultra-low write latency
  • highly available.
Data loss window if primary crashes before secondaries sync.
Semi-SynchronousPrimary blocks until at least one secondary acknowledges the write.Protects against single node failures without blocking on all secondaries.Moderate latency penalty.
  • Production clusters use Raft or Paxos for leader election β€” Bully is textbook-only and rarely deployed at scale:
Consensus AlgorithmLiveness HeartbeatSafety Validation
Raft ConsensusStrict leader heartbeats prevent follower election timeouts.Followers only vote for candidates with logs at least as up-to-date as theirs.
Paxos / Multi-PaxosProposers retry with higher ballot numbers on timeout.Quorum agreement on a single chosen value per slot.

In the room

State your RPO out loud: synchronous replication means RPO β‰ˆ 0 but higher write latency; async is faster but readers may lag. If they ask about split-brain, mention quorum and fencing.

6. Strategic Tradeoffs

Read scaling and write durability trade off β€” we name both sides:

BenefitCost
Synchronous Data Safety (guarantees secondaries store identical blocks before client return, ensuring zero lost updates)Write Availability Risk (if a secondary is slow or down, synchronous replication can block all writes)
Asynchronous Write Performance (acknowledges client writes instantly, keeping request loops fast and available)Replication Lag & Stale Reads (secondaries lag behind master, resulting in temporary read staleness across regions)

7. Failure / Bottleneck Awareness

Failover interviews pivot to RPO/RTO, split-brain, and flapping β€” we bring these up first:

⏱️ RPO & RTO Framing

RPO (Recovery Point Objective): How much data you can lose in a failure. Synchronous replication β†’ RPO β‰ˆ 0; async replication β†’ RPO equals unreplicated WAL (seconds to minutes).

RTO (Recovery Time Objective): How fast you must restore service. Leader election + DNS/proxy cutover + replica catch-up set the floor β€” state why your replication mode matches the product SLA.

πŸ“‰ Replica Lag & Stale Reads

Problem: Async replicas trail the primary by replication lag. Routing reads to lagging secondaries returns stale balances, inventory counts, or session state.

Mitigation: Monitor seconds_behind_master, route critical reads to the primary or a caught-up replica, or use read-your-writes routing after mutations.

🌩️ Split-Brain (Dual Masters)

Problem: After a partition, the minority side may promote a new leader while the old leader still accepts writes β€” both paths corrupt shared state.

Mitigation: Leaders must hold quorum (N/2 + 1). A leader in a minority partition steps down when it cannot reach a majority.

☠️ Flapping Failover

Problem: Aggressive heartbeat timeouts trigger elections on brief network blips, demoting a recovering primary and causing unnecessary churn.

Mitigation: Use conservative failover thresholds, jittered election timeouts, and warmup before returning a node to the write path.

8. Common HLD Usage

These production patterns show replication and election in the wild:

  • Relational DB Failovers (PostgreSQL/MySQL): High Availability setups utilize ZooKeeper or Consul to monitor database nodes, executing DNS updates to point traffic to the newly promoted replica.
  • Kafka partition Leadership (KRaft): Uses integrated Raft quorum metadata to coordinate partition leader promotions inside broker clusters without ZooKeeper external dependencies.

9. Decision Signals

Think replication when HA SLAs or read scaling force multiple database nodes:

🎯 Think Replication & Election When:
  • You are designing systems with extreme 99.999% high availability targets (SLA).
  • You must scale relational database read volumes globally (using Asynchronous secondary replicas).
  • You are architecting custom distributed state coordinators where single-active master states must be guaranteed.

11. Deep Dive (Optional)

Consensus Leader Election Mechanics (Raft Protocol)

In a Raft-based consensus cluster (e.g. etcd, Consul), leader election operates inside a strict logical state machine:

  1. Nodes operate in one of three states: **Leader**, **Follower**, or **Candidate**.
  2. The Leader sends periodic **Heartbeats** (AppendEntries commands with zero data) to followers to maintain authority.
  3. If a Follower stops receiving heartbeats for a randomized **Election Timeout** (typically 150ms-300ms), it transitions to a **Candidate** state.
  4. The Candidate increments the logical **Term** counter, votes for itself, and sends `RequestVote` RPC calls to all other nodes.
  5. Other nodes grant their vote if the candidate's logical Term is higher and its logged transactions index is at least as up-to-date as theirs.
  6. If the Candidate receives votes from a majority quorum (N/2 + 1 nodes), it transitions to **Leader** and immediately begins broadcasting heartbeats.

The randomized election timeout window is critical to prevent split votes where multiple candidates divide the quorum equally.

πŸ’¬Review

Help Us Improve

How helpful was this walkthrough?

Click a star to rate. We actively use this feedback to refine and update our system design content.

Placeholder
Optional but highly appreciated!

Discussion

Share your thoughts, ask questions, or help others.

Loading comments...