How Discord Solved the “@everyone” Problem
One @everyone ping sent millions of identical reads to one hot partition. How a Rust service that coalesces requests took latency from 150 ms to 15 ms.
If you’ve ever been part of a huge Discord server, you already know the chaos that follows when someone types @everyone. Your phone buzzes, the server explodes with activity, and tens of thousands of users check the message at the same time.
Now imagine the same chaos happening inside Discord’s backend.
This article explains how Discord almost melted its own database because of @everyone, and how it fixed the problem with a smart architectural pattern written in Rust.
Let’s break it down.
The Original Architecture: Partitioning by Channel#
Discord stored messages by channel_id. Simple, logical and efficient, until a massive channel triggers a storm.
- One message with @everyone
- Hundreds of thousands or millions of users open the app
- All of them request the exact same data row
- All requests go to the same database partition
This created a hot partition: one unlucky database node got hammered by millions of identical queries.
And because Discord’s database used a quorum-based system:
- If one node slowed down,
- The entire cluster slowed down.
Latency spiked from 15 ms to around 150 ms, a 10x slowdown.
At Discord’s scale, that’s unacceptable.
The Real Problem#
The message data for @everyone was the same for every user, but the system treated each request separately.
Imagine a million people knocking on the same door, asking the same question. The door won’t last long.
What Discord needed was:
- One fetch
- One worker
- A shared result for everyone
The Fix: A Rust-Based Data Service Layer#
Discord introduced a new layer between the API and the database: a data service written in Rust.
Why Rust?
- Extremely fast
- Memory-safe
- Great for high-performance concurrent systems
This new layer acted like a smart traffic controller.
How It Works#
Step 1: The first user triggers the @everyone request#
A worker is created inside the data service to fetch the message once.
Step 2: All other users “subscribe”#
Instead of hitting the database directly, their requests attach to that same worker.
No duplicate queries. No unnecessary load.
Step 3: The worker completes#
Once the worker has the message data:
- It immediately returns the result to every subscriber
- No repeated DB lookups
- No hot partitions
This is basically request deduplication + fan-out, done at extreme scale.
The Results#
After this architectural change:
- Millions of redundant DB queries vanished
- The hot partition disappeared
- The system stabilized under massive loads
- Latency dropped from 150 ms → ~15 ms
In other words: @everyone stopped being a backend nightmare.
Conclusion#
Discord didn’t fix the problem by throwing more servers or hardware at it. It solved it by understanding the root cause and designing a clean, intelligent architectural fix.
This pattern of one worker, many subscribers is useful in any system where:
- Many users request the same resource
- Data doesn’t change instantly
- You need to protect your database from repeated hits
Discord just implemented it at a scale where millions of people might request the same thing in the same second.
Smart engineering beats brute force every time.
Filed under hot-partitions, request-coalescing, databases
Written by Dhananjay Aggarwal