1. What It Is
RAID protects a single machine's disks from hardware failure — not from fires or accidental deletes. In cloud interviews we mention it when discussing bare-metal DB hosts or explaining why S3 uses erasure coding instead.
What:
RAID (Redundant Array of Independent Disks) coordinates multiple physical drives into a single logical volume managed by a hardware or software controller.
Primary purpose:
Protect physical hosts against hardware hard drive failures and increase disk read/write throughput concurrency.
Usually used for:
Bare-metal server storage clusters, network-attached storage (NAS) arrays, and database hosts disk pools.
2. Core Mental Model
RAID protects a single machine's disks from hardware failure — not from fires, ransomware, or accidental deletes. In cloud interviews, EBS already replicates under the hood; mention RAID when discussing bare-metal DB hosts or explaining why S3 durability is erasure-coded across racks, not mirrored RAID 1.
📊 Striping (Performance)
RAID 0 splits data files across multiple disks sequentially. This doubles read/write concurrency, but guarantees data loss on single disk crash.
🪞 Mirroring (Redundancy)
RAID 1 duplicates the exact same bytes onto a secondary disk clone. Ideal for zero-latency data resilience.
🧮 XOR Parity (Efficiency)
RAID 5 writes logical mathematical XOR blocks across drives. If one drive fails, the controller recomputes missing bytes on-the-fly.
In the room
Don't propose RAID for a cloud-native design — EBS already replicates. RAID matters for on-prem DB servers or when explaining durability layers. Know RAID 0 (striping, no redundancy), RAID 1 (mirror), RAID 5/6 (parity) at a high level.
3. Why It Matters in HLD
RAID trades disks for redundancy — interviewers use it to discuss durability before distributed replication. Three lenses:
Needed When:
Configuring bare-metal database instances, sizing cloud storage IOPS bounds, or planning physical high-availability datacenters.
Avoids:
Catastrophic host downtimes from routine hardware disk crashes, database write performance bottlenecks, and single-drive capacity halts.
Optimizes For:
Disk write IOPS throughputs, single-host liveness SLAs, hardware capacity ratios, and physical storage costs.
4. Architecture & Data Flow
Walk RAID levels as interview comparison steps. Step 1 — RAID 0: striping only — fast, zero redundancy. Step 2 — RAID 1: mirror — 2x capacity cost, survives one disk. Step 3 — RAID 5: parity striping — survives one disk, rebuild penalty. Step 4 — RAID 10: mirror + stripe — common DB choice, survives disk failure with lower rebuild risk. Step 5 — Distributed step: note RAID protects one machine; cross-node replication protects the rack.
5. Key Characteristics
Capacity overhead, read/write performance, and fault tolerance vary by level — we compare:
- RAID level comparison — fault tolerance, I/O profile, and typical use:
| RAID Level | Fault Tolerance Limit | Read / Write Performance Impact | Primary Use Case |
|---|---|---|---|
| RAID 0 (Striping) | 0 Disks (Any disk failure crashes the entire partition). | Excellent Read & Write (concurrency across all disks). | Non-critical temporary caches, scratch files. |
| RAID 1 (Mirroring) | N - 1 Disks (Can survive until last mirrored replica fails). |
| OS Boot drives, critical single-database WAL partitions. |
| RAID 5 (Parity) | 1 Disk (Survives single disk crash via XOR parity blocks). |
| Large analytical read-heavy clusters, corporate NAS volumes. |
| RAID 10 (1+0 Striped Mirrors) | Up to N/2 Disks (1 disk per mirror pool). | Ultra-fast Read & Write (combines mirroring and striping). | High-throughput production database storage disks (SQL/NoSQL). |
Erasure coding (cloud-scale durability): RAID parity protects one disk in one server. Object stores (S3, GCS) split each object into k data fragments + m parity fragments (e.g., 8+3 Reed-Solomon) distributed across separate racks and AZs. You can lose any 3 fragments and still reconstruct — ~1.375x storage overhead vs RAID 1's 2x. Glacier and cold tiers rely on erasure coding, not RAID mirrors. In interviews: RAID for on-prem DB hosts; erasure coding for distributed blob durability (ties to concept #21 Storage Types and problem #87 S3).
In the room
Say "RAID is not backup" — it protects disk failure, not accidental delete or ransomware. Interviewers nod when you distinguish redundancy from backup.
6. Strategic Tradeoffs
Local redundancy is simpler than distributed replication but does not survive node loss — we state both:
| Benefit | Cost |
|---|---|
| Server Uptime Protection (survives hard drive hardware crashes without server reboots, maintaining SLAs) | Storage Capacity Penalty (mirroring and parity require 30-50% redundant capacity overheads) |
| Enhanced IOPS Speed (striping routes read/write queues across multiple spindles concurrently, bypassing single disk limits) | No Disaster Protection (does not shield against accidental file deletions, malware, or complete datacenter fires) |
7. Failure / Bottleneck Awareness
Rebuild storms, URE during rebuild, and RAID-not-backup confusion — we volunteer:
Problem: RAID mirrors or parity protect against disk failure, not user error. A deleted folder or ransomware infection replicates across all drives immediately.
Mitigation: Keep isolated, versioned backups in separate object storage with retention locks — independent of the RAID array.
Problem: Each RAID 5 write reads old data and parity, recomputes parity, then writes both blocks — four I/O operations per logical write. Write-heavy databases suffer.
Mitigation: Avoid RAID 5 for transactional write workloads. Prefer RAID 10 (striped mirrors) for concurrent read/write performance without parity recompute cost.
8. Common HLD Usage
Database servers and NAS appliances cite RAID levels in HLD discussions:
- Transactional Database Servers (RAID 10): Enterprise database clusters combine mirroring safety and striped performance to maximize disk transaction speeds.
- Glacier Cold Object Storage (Erasure Coding): Modern cloud stores bypass physical RAID limits entirely by using distributed Erasure Coding, splitting user data into N data chunks + M parity fragments distributed across separate datacenter regions.
9. Decision Signals
Mention RAID when discussing single-node durability before jumping to distributed replicas:
- You are designing physical database host infrastructure requiring stable server-level disk resilience.
- Your disk write/read requirements outgrow the performance limitations of a single SSD drive.
- You want to explain RAID 5 write penalties or XOR parity when justifying RAID 10 for OLTP workloads.
- The design is fully managed cloud (RDS, DynamoDB, S3) — say "provider handles disk replication" and pivot to erasure coding or cross-AZ replication instead.
- The problem is application-level durability (WAL, Kafka ISR) — RAID is below that abstraction.
11. Deep Dive (Optional)
The XOR Math behind RAID 5
To secure data with minimal storage overhead, RAID 5 leverages the logical **Exclusive OR (XOR)** operator:
Data Block 1 = 10101010
Data Block 2 = 11110000
Parity Block = Block 1 ⊕ Block 2
= 01011010If Disk A crashes (losing Data Block 1), the controller recomputes it dynamically by executing the XOR operation across the surviving drive and the parity block:
Block 1 = Block 2 ⊕ Parity Block
= 11110000 ⊕ 01011010
= 10101010 (Perfect recovery!)This mathematically restores missing bytes on-the-fly, sacrificing slight CPU computation cycles to achieve substantial hardware economy.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.