Designing a Scalable URL Shortener Like TinyURL: A Complete System Design Breakdown
APIs, Base62 codes, Zookeeper-assigned ID ranges, Cassandra blind writes and Redis caching: how a TinyURL-style service scales without collisions.
A production-grade URL shortening service looks simple on the surface. A user enters a long URL, the system returns a short one, and clicking it redirects you to the original page. Behind that simplicity sits a highly optimized distributed system that processes billions of links and handles massive spikes in traffic.
This article breaks down the core system design of a TinyURL-like service. It is based on a mock system design interview in which a candidate designs a high-throughput link shortening platform. The goal is to understand the real-world architecture patterns, trade-offs and scaling strategies that make such services reliable at global scale.
1. Problem Overview#
A URL shortener must provide two core functions:
- Accept a long URL and generate a unique short URL.
- Redirect users extremely fast when they click a short link.
The system must scale horizontally, avoid collisions in short code generation, cache aggressively, and support billions of records while keeping latency low.
2. Designing the Core APIs#
A clean API surface is essential for scalability and for easy integration with clients.

Shortening API#
HTTP Method: POST
Endpoint: /shorten
Request Body:
Response:
The server receives the long URL, generates a unique ID, stores the mapping, and returns the shortened version.
Redirect API#
HTTP Method: GET
Endpoint: /{short_code}
Process flow:
- Read the short code from the URL.
- Look up the long URL in the cache or database.
- Return an HTTP 301 Permanent Redirect to the original link.
The 301 redirect lets browsers and crawlers cache the destination, which reduces repeated load on the backend.

3. How Short Codes Are Generated#
Generating short, unique codes is central to the system. The chosen method is Base62 encoding.

What is Base62?#
It uses 62 characters:
- 26 lowercase letters [a-z]
- 26 uppercase letters [A-Z]
- 10 digits [0–9]
That gives 62 possibilities for each character.
Why 7 characters?#
Total possible combinations: 62⁷ ≈ 3.5 trillion
That supports several decades of URL creation at a scale of billions per year.
4. Horizontal Scaling and the Collision Challenge#
With one application server generating IDs, collisions are easy to avoid. But real systems need multiple servers behind a load balancer.

If 10 servers generate short codes at the same time, the chance of accidentally generating the same code rises unless they are coordinated.
Problems with the naive approach#
- A central lock on the database slows down writes.
- Checking for collisions on every write is expensive.
- Generating random IDs on multiple servers risks overlaps.
We need a deterministic, distributed and collision-free strategy.
5. The Distributed ID Solution Using Zookeeper#
A highly scalable approach is to use Zookeeper as a coordination service that assigns non-overlapping numeric ranges to each application server.
Example:
- Server 1 gets 0 to 1,000,000
- Server 2 gets 1,000,001 to 2,000,000
- Server 3 gets 2,000,001 to 3,000,000
Every server keeps an internal counter. When a request comes in:
- The counter increments.
- The integer is converted to Base62.
- The resulting code is unique, because its underlying numeric range never overlaps with another server’s.

Advantages#
- Zero collisions
- No locking
- No database uniqueness checks
- Extremely fast generation
- Perfect for high-throughput systems
This is the same principle behind Twitter Snowflake-style architectures.
6. Database Strategy: Cassandra for High Velocity Writes#

The mapping between short code and long URL must be stored durably and be retrievable in milliseconds.
Technology choice: Apache Cassandra
Reasons:
- Highly scalable distributed NoSQL database
- Optimized for very fast writes
- High availability with multi-data-center replication
- Ideal for massive key-value workloads
The schema is simple:
Partition key: short_code Column: long_url
Because the ID generation system guarantees uniqueness, database writes are blind inserts with no prior checks. This removes expensive read-before-write operations.
7. Handling Read Heavy Traffic With Redis Caching#
In real systems, reads are far more frequent than writes. A popular link may get millions of clicks a day.
Read flow#
- A user clicks the short link.
- The application server checks Redis.
- On a cache hit, it returns the long URL in microseconds.
- On a cache miss, it:
- fetches from Cassandra
- populates Redis
- redirects the user
Benefits of Redis#
- In-memory store
- Millisecond latency
- Easy to cluster
- Reduces database load by 80 to 95 percent
Since redirects need to be extremely fast, Redis gives the system a large performance boost.
8. Architectural Diagram#

9. Handling Scalability and Fault Tolerance#
This design supports:
- Horizontal scaling of application servers
- Independent scaling of the cache and database layers
- Fault tolerance through Cassandra replication
- Crash safety through Zookeeper’s consistency guarantees
Components can fail independently without taking the whole service down.
Key scalability insights:
- Stateless app servers are easy to scale out.
- Redis clustering handles massive read throughput.
- The Cassandra ring supports petabyte-scale storage.
- Zookeeper prevents race conditions in distributed ID generation.
This architecture can handle billions of URLs and millions of requests per second.
10. Learnings and Why This Matters#
Designing systems like URL shorteners teaches foundational concepts used across major tech companies.
Key takeaways:
- Distributed ID generation patterns
- Consistent hashing and partitioning
- The cache-aside strategy
- NoSQL store design
- Load balancing and horizontal scaling
- Avoiding global locks
- Keeping data consistent at scale
URL shorteners look tiny, but the ecosystem behind them is a masterclass in high-performance distributed systems.
Conclusion#
A robust URL shortening service needs more than string manipulation. It requires:
- Carefully structured APIs
- Reliable range-based ID generation
- Extremely fast caching
- Scalable storage
- Thoughtful consistency and availability strategies
These patterns carry over to real-world systems like social networks, authentication services and large-scale logging platforms.
If you are preparing for system design interviews or building large-scale backend systems, this architecture is a must-study example. For the full interview walkthrough, from requirements to trade-offs, see Design a URL Shortener.
Filed under url-shortener, system-design-interview, distributed-ids
Written by Dhananjay Aggarwal
