Designing a Scalable URL Shortener Like TinyURL: A Complete System Design Breakdown

APIs, Base62 codes, Zookeeper-assigned ID ranges, Cassandra blind writes and Redis caching: how a TinyURL-style service scales without collisions.

Dhananjay Aggarwal, · 7 min read
Share
Summarize with AI
A glowing chain-link icon above the title Designing a Scalable URL Shortener Like TinyURL: A Complete System Design Breakdown.

A production-grade URL shortening service looks simple on the surface. A user enters a long URL, the system returns a short one, and clicking it redirects you to the original page. Behind that simplicity sits a highly optimized distributed system that processes billions of links and handles massive spikes in traffic.

This article breaks down the core system design of a TinyURL-like service. It is based on a mock system design interview in which a candidate designs a high-throughput link shortening platform. The goal is to understand the real-world architecture patterns, trade-offs and scaling strategies that make such services reliable at global scale.

1. Problem Overview#

A URL shortener must provide two core functions:

  1. Accept a long URL and generate a unique short URL.
  2. Redirect users extremely fast when they click a short link.

The system must scale horizontally, avoid collisions in short code generation, cache aggressively, and support billions of records while keeping latency low.

2. Designing the Core APIs#

A clean API surface is essential for scalability and for easy integration with clients.

POST /shorten stores the mapping from long URL to short URL and returns the short URL to the user.

Shortening API#

HTTP Method: POST Endpoint: /shorten Request Body:

JSON
{
  "long_url": "https://example.com/abc..."
}

Response:

JSON
{
  "short_url": "https://tinyurl.com/aB3xK7p"
}

The server receives the long URL, generates a unique ID, stores the mapping, and returns the shortened version.

Redirect API#

HTTP Method: GET Endpoint: /{short_code}

Process flow:

  1. Read the short code from the URL.
  2. Look up the long URL in the cache or database.
  3. Return an HTTP 301 Permanent Redirect to the original link.

The 301 redirect lets browsers and crawlers cache the destination, which reduces repeated load on the backend.

A user opens a short URL, GET /{code} looks up the mapping, and the service answers with an HTTP 301 to the long URL.

3. How Short Codes Are Generated#

Generating short, unique codes is central to the system. The chosen method is Base62 encoding.

Base62 character set chart: 26 uppercase letters A-Z, 26 lowercase letters a-z and 10 digits 0-9, 62 characters in total.

What is Base62?#

It uses 62 characters:

  • 26 lowercase letters [a-z]
  • 26 uppercase letters [A-Z]
  • 10 digits [0–9]

That gives 62 possibilities for each character.

Why 7 characters?#

Total possible combinations: 62⁷ ≈ 3.5 trillion

That supports several decades of URL creation at a scale of billions per year.

4. Horizontal Scaling and the Collision Challenge#

With one application server generating IDs, collisions are easy to avoid. But real systems need multiple servers behind a load balancer.

High traffic of 1K URLs per second goes through a load balancer to three servers that all write to one database.

If 10 servers generate short codes at the same time, the chance of accidentally generating the same code rises unless they are coordinated.

Problems with the naive approach#

  • A central lock on the database slows down writes.
  • Checking for collisions on every write is expensive.
  • Generating random IDs on multiple servers risks overlaps.

We need a deterministic, distributed and collision-free strategy.

5. The Distributed ID Solution Using Zookeeper#

A highly scalable approach is to use Zookeeper as a coordination service that assigns non-overlapping numeric ranges to each application server.

Example:

  • Server 1 gets 0 to 1,000,000
  • Server 2 gets 1,000,001 to 2,000,000
  • Server 3 gets 2,000,001 to 3,000,000

Every server keeps an internal counter. When a request comes in:

  1. The counter increments.
  2. The integer is converted to Base62.
  3. The resulting code is unique, because its underlying numeric range never overlaps with another server’s.
The same load-balanced setup, with each of the three servers now owning its own ID range before writing to the database.

Advantages#

  • Zero collisions
  • No locking
  • No database uniqueness checks
  • Extremely fast generation
  • Perfect for high-throughput systems

This is the same principle behind Twitter Snowflake-style architectures.

6. Database Strategy: Cassandra for High Velocity Writes#

The write path: user, load balancer, an app server with range 1M to 2M, Base62 conversion to aB3xK7p, and an INSERT with no uniqueness check.

The mapping between short code and long URL must be stored durably and be retrievable in milliseconds.

Technology choice: Apache Cassandra

Reasons:

  • Highly scalable distributed NoSQL database
  • Optimized for very fast writes
  • High availability with multi-data-center replication
  • Ideal for massive key-value workloads

The schema is simple:

Partition key: short_code Column: long_url

Because the ID generation system guarantees uniqueness, database writes are blind inserts with no prior checks. This removes expensive read-before-write operations.

7. Handling Read Heavy Traffic With Redis Caching#

In real systems, reads are far more frequent than writes. A popular link may get millions of clicks a day.

Read flow#

  1. A user clicks the short link.
  2. The application server checks Redis.
  3. On a cache hit, it returns the long URL in microseconds.
  4. On a cache miss, it:
    • fetches from Cassandra
    • populates Redis
    • redirects the user

Benefits of Redis#

  • In-memory store
  • Millisecond latency
  • Easy to cluster
  • Reduces database load by 80 to 95 percent

Since redirects need to be extremely fast, Redis gives the system a large performance boost.

8. Architectural Diagram#

Full architecture: users hit a load balancer in front of three app servers that get ID ranges from Zookeeper, read through a Redis cache layer, and store mappings in a Cassandra database layer.

9. Handling Scalability and Fault Tolerance#

This design supports:

  • Horizontal scaling of application servers
  • Independent scaling of the cache and database layers
  • Fault tolerance through Cassandra replication
  • Crash safety through Zookeeper’s consistency guarantees

Components can fail independently without taking the whole service down.

Key scalability insights:

  • Stateless app servers are easy to scale out.
  • Redis clustering handles massive read throughput.
  • The Cassandra ring supports petabyte-scale storage.
  • Zookeeper prevents race conditions in distributed ID generation.

This architecture can handle billions of URLs and millions of requests per second.

10. Learnings and Why This Matters#

Designing systems like URL shorteners teaches foundational concepts used across major tech companies.

Key takeaways:

  • Distributed ID generation patterns
  • Consistent hashing and partitioning
  • The cache-aside strategy
  • NoSQL store design
  • Load balancing and horizontal scaling
  • Avoiding global locks
  • Keeping data consistent at scale

URL shorteners look tiny, but the ecosystem behind them is a masterclass in high-performance distributed systems.

Conclusion#

A robust URL shortening service needs more than string manipulation. It requires:

  • Carefully structured APIs
  • Reliable range-based ID generation
  • Extremely fast caching
  • Scalable storage
  • Thoughtful consistency and availability strategies

These patterns carry over to real-world systems like social networks, authentication services and large-scale logging platforms.

If you are preparing for system design interviews or building large-scale backend systems, this architecture is a must-study example. For the full interview walkthrough, from requirements to trade-offs, see Design a URL Shortener.

Filed under url-shortener, system-design-interview, distributed-ids

Was this post useful?
Share
Summarize with AI
Prefer IntervueClub on GoogleShow our posts more often in Top Stories

Written by Dhananjay Aggarwal

Shard by user or by time? Work it out with the write rateOct 4, 2026 · 4 min readEstimating QPS from daily users in three stepsSep 28, 2026 · 2 min read

All posts