1. What It Is
CAP comes up the moment you draw more than one database box. During a network partition you must choose consistency or availability β and PACELC reminds us that even without a partition, latency and consistency still trade off.
What:
During a network partition in a distributed system, you can guarantee Consistency (CP) OR Availability (AP), but not both.
Primary purpose:
Guide your trade-offs around database scaling, distributed locking, and partition resilience during network splits.
Usually used for:
Designing microservices state stores, selecting distributed datastores, and defining consistency SLAs.
2. Core Mental Model
Partitions are inevitable β your design must state what happens to reads and writes when the network splits:
β‘ PACELC Extension
CAP only applies when there is a partition (P). Else (E), you must balance Latency (L) against Consistency (C).
π€ Quorum overlap
Tune consistency dynamically: if your read node pool overlaps your write pool (R + W > N), you read the latest value.
π°οΈ Vector Clock Timestamps
In eventual AP networks, use logical vector clocks to trace dependencies and resolve concurrent write conflicts.
In the room
Don't recite "pick two of three" and stop. Interviewers want to hear which operations are CP vs AP in your design β a shopping cart can be AP while the payment ledger stays CP. Mention PACELC if they push on normal-operation latency.
3. Why It Matters in HLD
CAP is not a trivia question β it forces us to state what happens to reads and writes during a partition. We use three lenses:
Needed When:
Whiteboarding stateful multi-datacenter clusters, distributed locks, or high-throughput shopping databases.
Avoids:
Double-spend bank records, microservices sync deadlocks, and stale data propagation in critical payment paths.
Optimizes For:
Partition fault survivability, predictable read latencies, write performance throughput, and data convergence safety.
4. Architecture & Data Flow
Walk quorum overlap as interview steps. Step 1 β Normal path: client writes to W nodes, reads from R nodes; when R + W > N, reads see the latest write. Step 2 β Partition hits: network splits replicas into minority and majority camps. Step 3 β CP choice: minority-side writes fail; clients get errors but no stale commits. Step 4 β AP choice: both sides accept writes; reconciliation happens after heal. Step 5 β Name per operation: cart is AP, payment ledger is CP β mixed models in one system.
CAP Under Network Partition
When the network splits, replicas cannot talk β the system must choose whether minority-side writes fail (CP) or continue with later reconciliation (AP):
5. Key Characteristics
We place the design on the CAP and consistency spectrum β weakest level that satisfies the product invariant:
- During a partition, CP systems reject minority writes; AP systems accept and reconcile later:
| Option Profile | Partition Behavior | Primary Use Case |
|---|---|---|
| CP (Consistency + Partition) | Reject writes / returns error to block stale updates. | Financial ledgers, key coordination registries (Consul, ZooKeeper). |
| AP (Availability + Partition) |
| Social timeline updates, chat histories, shopping carts (DynamoDB). |
- Consistency models sit on a spectrum from linearizable to eventual β name the weakest level that satisfies your invariant:
| Consistency Model | Behavioral Contract | Performance Overhead |
|---|---|---|
| Linearizable (Strong) | All reads return the absolute latest committed write globally. | Extremely high (requires strict synchronization locks across nodes) |
| Eventual Consistency |
| Minimal (zero blocking, asynchronous updates) |
| Causal Consistency | Operations that are causally related are seen in correct order. | Moderate (requires logical vector clocks tracking dependencies) |
In the room
Do not recite "pick two of three" and stop. Say which operations are CP vs AP in your design β that single move separates memorization from engineering judgment.
6. Strategic Tradeoffs
Strong consistency and high availability pull apart during partitions. We state both sides:
| Benefit | Cost |
|---|---|
| Linearizable Consistency (guarantees a single global state, entirely eliminating race conditions or out-of-order logs) | High Latency / Downtime (during network partition splits, CP systems reject writes or block queries, stalling users) |
| High Availability (AP) (system accepts writes on any isolated node partition, ensuring zero downtime for customers) | Stale Reads & Write Conflicts (clients can read stale states, and conflicting writes require complex reconciliation) |
7. Failure / Bottleneck Awareness
CAP misunderstandings are common interview traps β we dispel them and name real failure modes:
Myth: "Pick CP or AP once for the whole system." Reality: CAP applies per operation during a partition β a shopping cart can be AP while the payment ledger stays CP.
Myth: "NoSQL is always AP; SQL is always CP." Reality: PostgreSQL with synchronous replicas is CP; Cassandra with QUORUM reads can be tuned toward stronger consistency.
Myth: "CAP means you abandon consistency or availability forever." Reality: The tradeoff binds only while the partition persists β once the network heals, both can return.
Problem: A partition can leave two halves each believing they are the leader, both accepting writes and diverging state.
Mitigation: Require quorum (majority) for leadership and writes so only the larger partition stays writable.
Problem: Resolving conflicts by wall-clock timestamp can drop newer writes when clocks drift between nodes.
Mitigation: Use vector clocks or Lamport timestamps to preserve causal order instead of raw system time.
8. Common HLD Usage
Production systems mix CP and AP by domain β we cite these patterns to show nuance:
Real designs mix consistency models by domain β coordination stores stay CP while user-facing writes may accept eventual convergence:
- Distributed Locking (CP): Systems like ZooKeeper or Raft consensus engines reject updates if quorum fails, protecting critical single-owner locks.
- Dynamic Shopping Carts (AP): Retail carts (e.g. Amazon Dynamo) accept all offline additions to prioritize sales, resolving item counts later at checkout via CRDT merging.
9. Decision Signals
Reach for CAP framing when multi-region sync or consensus backbones enter the design:
- You are designing multi-region database sync topologies.
- You must choose between strongly-consistent consensus backbones (Raft/ZooKeeper) and highly-available Cassandra clusters.
- You are architecting system configurations where data conflicts (like concurrent seat reservations) are catastrophic.
11. Deep Dive (Optional)
In interviews, default to eventual consistency unless the problem involves money, inventory counts, or booking finite resources (seats, tickets). Many systems mix models: product copy can lag; stock and payments cannot.
Sloppy Quorums vs Strict Quorums
To maximize availability, AP wide-column databases (Cassandra, DynamoDB) deploy **Sloppy Quorums** combined with **Hinted Handoffs**:
- Under normal operations, writes map to strict replica nodes on the hash ring.
- If primary replica nodes are down during a partition, the write is accepted by a neighboring node on the ring.
- The neighboring node saves a **hint** (metadata block) in its local database storage.
- Once the network partition heals and the primary replica returns online, the neighbor hands off the updates asynchronously.
This boosts system write availability under high partition degradation, at the cost of temporary read staleness.
PACELC: Beyond the Binary CAP Choice
CAP only describes behavior during a network partition. PACELC extends the lens to normal operation: if there is a Partition, choose A or C; Else (no partition), choose Latency or Consistency. DynamoDB and Cassandra bias toward low latency even when the network is healthy β reads may return slightly stale replicas to avoid cross-region round trips.
Consistency Levels on a Spectrum
Interviewers rarely want a single word answer. Place your design on this ladder and justify the trade-off:
- Strong / linearizable: Reads reflect the latest successful write globally. Needed for bank balances, inventory locks. Cost: higher latency, lower availability under partition.
- Causal: If operation A happened before B, all readers see that order. Good for social comments and collaborative docs without paying for global linearizability on every read.
- Read-your-writes: A user always sees their own updates immediately; other users may lag. Session stickiness or routing reads to the write leader achieves this cheaply.
- Eventual: Replicas converge after writes stop; acceptable for like counts, view counters, and CDN cache propagation when product tolerates seconds of skew.
Strong signal: name the weakest consistency level that still satisfies the product invariant β not the strongest level you know how to spell.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.