1. What It Is
Picking storage isn't just "S3 or database" β block, file, and object stores solve different problems. We match the abstraction to access pattern: random I/O, shared filesystem, or immutable blobs.
What:
Comparing Block (raw disk partitions), File (hierarchical shared folders), and Object (flat metadata buckets) storage architectures.
Primary purpose:
Matching workload I/O profiles (low-latency writes vs massive immutable files) to the optimal storage engine to balance performance, cost, and scale.
Usually used for:
Database transactional volumes, shared legacy app server mounts, and large media files distributions.
2. Core Mental Model
Pick storage by access pattern, not brand names: random small writes β block; shared folder tree across pods β NFS/EFS; large immutable blobs at web scale β object. Problem #87 Blob Storage (S3) is the object-store HLD; concept #24 GFS/HDFS covers cluster file systems for analytics.
πΎ Raw Block Sectors
Like a blank slate disk volume. Operating systems format Block drives with native filesystems (ext4) to host databases.
π Hierarchical Files
Organized in nested folders. Shared File systems (NFS) enable multiple hosts to read/write shared configuration files.
πͺ£ Flat HTTP Objects
No nested folders. Objects represent key-value states accessed via HTTP APIs. Ideal for holding static media assets.
In the room
For most web apps, object storage (S3) handles media and backups; block storage backs databases. File storage (NFS/EFS) appears when legacy apps need a shared POSIX filesystem. Don't put a database on object storage.
3. Why It Matters in HLD
Storage tier choice follows access pattern β block for databases, file for shared POSIX, object for blobs. Three lenses:
Needed When:
Sizing databases, designing media ingress systems, configuring shared application assets, or calculating monthly hosting budgets.
Avoids:
Database CPU/disk stalls due to network storage latency spikes, high cloud costs from misallocated block volumes, and file corruptions.
Optimizes For:
Storage cost efficiency, data throughput speeds, random write capacities, and absolute storage scale limits.
4. Architecture & Data Flow
Walk data placement as interview steps. Step 1 β Hot structured: block storage (EBS) attached to DB servers for low-latency random I/O. Step 2 β Shared files: NFS/EFS when multiple nodes need POSIX paths. Step 3 β Blobs: object store (S3) for images, video, backups β HTTP GET/PUT, 11-nines durability. Step 4 β Lifecycle: tier cold data to Glacier; state retrieval latency trade-off.
In the room
Say S3 plus presigned multipart upload for large files β not "store video in PostgreSQL BLOB columns." That redirect shows storage maturity.
5. Key Characteristics
Block vs file vs object differ on interface, consistency, and cost β we compare:
- Storage type comparison β access model, latency, and scale limits:
| Storage Dimension | Access Interface | Access Latency Bounds | Capacity Scalability |
|---|---|---|---|
| Block Storage | Sector volumes over FC/iSCSI. | Sub-millisecond (ultra-fast raw disk access). | Fixed-capacity volumes (requires manual volume extension). |
| File Storage | Shared directory paths over NFS/SMB. | Moderate (network lookup overheads). | Elastic up to file-server storage limits. |
| Object Storage | REST HTTP API calls (GET/PUT). | High (HTTP handshake roundtrips). | Strictly infinite capacity scaling. |
- File/NFS tier in practice β NFS (or cloud EFS/Azure Files) suits legacy apps expecting
/mnt/shared/config, ML training sets read by many GPU nodes, and video transcode pipelines where workers need hierarchical paths. Trade-offs: metadata latency (everyopen()is a network round-trip), weak cross-region performance, and file-lock semantics that differ from local ext4. Prefer object storage for internet-scale uploads; NFS when POSIX directory semantics are required. - Multipart object upload β for files >100 MB, clients split into parts (e.g., 8 MB each), upload in parallel via presigned URLs, then call
CompleteMultipartUpload. Failed parts retry individually without restarting the whole file. Mention this in any large-blob interview (#15 Video Streaming, #25 Dropbox, #87 S3).
6. Strategic Tradeoffs
Durability and scale trade latency and semantic richness β we articulate both:
| Benefit | Cost |
|---|---|
| Block low latency (interacts directly with raw sectors, bypassing OS filesystem abstractions for maximum database performance) | No Multi-Node Writes (cannot be attached to multiple active servers concurrently without immediate block corruption) |
| Object infinite scale (flat directory structure maps keys to values, scaling to Exabytes without sizing updates) |
|
7. Failure / Bottleneck Awareness
Wrong tier for the workload, small-file overhead in object stores β we name pitfalls:
Problem: Attaching one EBS volume to two active instances for concurrent writes. Most filesystems assume exclusive ownership; dual writers corrupt metadata and destroy the partition.
Mitigation: Do not share raw block volumes across writers. Use a shared file system (NFS, EFS) with coordinated locking, or route through object storage APIs.
Problem: Object stores have a flat key namespace with no SQL or directory queries. Scanning billions of objects to find "uploaded_by = user_123" is slow and expensive.
Mitigation: Keep a secondary index in a database (PostgreSQL, DynamoDB) mapping business keys to object URLs; treat the bucket as blob storage only.
8. Common HLD Usage
Video, backups, and user uploads map to object; databases map to block:
| Production Workload | Selected Storage Tier | Architectural Rationale |
|---|---|---|
| PostgreSQL Data Directory | Block Storage (AWS EBS GP3) | High-velocity random transaction updates require raw block sector access and stable, provisioned IOPS throughput. |
| Shared Config / Media Processing Staging | File Storage (NFS / AWS EFS) | Multiple transcode workers mount the same directory tree to read source files and write intermediate chunks β hierarchical paths matter more than HTTP key semantics. |
| YouTube User Video Uploads | Object Storage (AWS S3) + multipart upload |
|
9. Decision Signals
Pick object storage when payloads are large blobs with HTTP access; block when random I/O matters:
- You are designing databases requiring high-concurrency random block updates.
- You must scale media asset pipelines (user avatars, videos, invoice PDFs) to handle infinite data growth at minimal hosting costs.
- You need to mount a shared directory path across multiple legacy monolith application containers.
11. Deep Dive (Optional)
S3 Object Lifecycle Policies (Cloud Budget Economy)
For massive, high-volume analytics systems, static assets and log files accumulate rapidly. Retaining all raw files inside primary High-Performance Object storage tiers forever wastes immense budget resources.
Solution: tiered lifecycle transitions
- Active Tier (0 - 30 days): Store new logs inside standard S3 buckets for instant sub-second retrieval.
- Infrequent Access Tier (30 - 90 days): Automatically transition logs to **S3-IA** (Infrequent Access) to cut hosting costs by 50% (introduces small retrieval fees).
- Glacier Archive Tier (90 - 365 days): Transition logs to **S3 Glacier Flexible** for long-term backups (reduces storage costs by 80%, retrieval takes 1-5 hours).
- Permanent Deletion (365+ days): Configure strict deletion policies to purge expired log files, keeping storage footprint bounded.
Review
How helpful was this walkthrough?
Click a star to rate. We actively use this feedback to refine and update our system design content.
Discussion
Share your thoughts, ask questions, or help others.