Day 11 B-Building blocks 2026-10-05 ← All lessons

Design a notification system: queues, retries, fan-out

Ten million pushes, five million emails, and one million texts leave one system every day, and none may be lost. This lesson builds that system: scope the load, meet each delivery channel, break the single server into queues and workers, then add the retry and guardrails that make it production.

16M / dayNotifications sent per day
10M pushMobile push per day
5M emailEmails per day
1M SMSSMS messages per day

Key points

Outcomes

01Derive the per-second send load from 16M notifications a day and explain why one server cannot absorb it.
02Trace one notification from a service API call through queue, worker, and third-party service to the device, including the retry loop.
03Argue why delivery can be at-least-once but never exactly-once, and predict what dedupe plus retry do about it.
04Choose per-type queues, opt-in checks, and rate limits when a provider fails or users feel spammed.
01

Scope first: 16 million a day, three channels


The interview opens with scope questions, and the answers fix four numbered requirements. (1) Three channels: mobile push, SMS, and email. (2) Soft real time: a user should get each notification as soon as possible, but a slight delay under heavy load is acceptable. (3) Volume: 10 million pushes, 1 million SMS, and 5 million emails per day. (4) Respect opt-out: a user who opts out receives nothing further. Triggers come from client applications or from server-side schedules, and supported devices span iOS, Android, and laptop or desktop.

Push per day
10M
Email per day
5M
SMS per day
1M
Push dominates the day at ten times SMS volume; email sits in the middle at five times.

Run the envelope live. 16 million a day divided by 86,400 seconds is about 185 sends per second on average: roughly 116 push, 58 email, and 12 SMS. That rate is only the floor. Each send also pays for rendering content, a database or cache lookup for contact info, and a round trip wait on a third-party service. One server absorbing 185 of those compound sends per second, with peaks above it, is the load that breaks the first design in the next section.

02

Each channel delivers differently


Each channel delivers through different hands. An iOS push needs three parts: a provider that builds the request with a device token and a payload, Apple Push Notification Service which relays it, and the device itself. The payload is a JSON dictionary with an alert title and body plus a badge count, as in Bob asking you to play chess. Android follows the same path through Firebase Cloud Messaging instead of APNs. SMS goes through commercial gateways such as Twilio or Nexmo. Email usually goes through services such as Sendgrid or Mailchimp rather than a self-built mail server, for better delivery rates and analytics.

Contact info lives in two simplified tables. The user table holds email addresses and phone numbers; the device table holds device tokens, and one user can own several devices, so one push can fan out to every device that user logged in on. Extensibility matters here because providers differ by market: FCM is unavailable in China, where alternatives such as Jpush and PushY take its place. A design that hardcodes one provider per channel cannot enter those markets.

Option A

Self-hosted delivery

  • You own the whole path and owe every retry, blocklist negotiation, and analytics pipeline to yourself
  • One provider outage or one new market means reworking delivery code, not swapping a plug-in
  • Spam and abuse controls are yours to build from nothing
Option B

Third-party services

  • Delivery rate, analytics, and scaling are the vendor's problem under a maintained API
  • A new provider plugs in beside the old one, which is how the China case is handled
  • You still own payload building, contact lookup, and the decision of what to send
03

From one server to queues and workers


The first design is one notification server between triggering services and the third-party services, and it fails in three numbered ways. (1) Single point of failure: one server down means no channel sends anything. (2) Hard to scale: database, cache, and per-channel processing cannot grow independently inside one box. (3) Performance bottleneck: HTML rendering plus waiting on third-party responses burns the one server fastest exactly at peak hours. Requirement (3) is what turns these from theory into an outage.

How to read: Step through the six stages in order, then follow the retry edge from Workers back to the queues and ask what happens when the third-party service is down.

Service NNotif serversPer-type queuesWorkers3rd-party svcUser device
  1. A triggering service calls the send API on the notification servers

  2. Servers validate addresses, fetch contact info from cache or DB, and enqueue

  3. The event waits in its own type queue: iOS, Android, SMS, or email

  4. A worker pulls the event and calls the matching third-party service

  5. On failure the event returns to the queue for retry instead of being lost

One event travels left to right through its own type queue; a failed send loops back to the queue for retry instead of being dropped.

The improved design moves the database and cache out of the server, runs many notification servers behind autoscaling, and puts a message queue between stages. Servers now do only four jobs: expose internal APIs to triggering services, validate emails and phone numbers, fetch rendering data from cache or DB, and enqueue events. Workers pull events and call third parties. Each notification type owns its queue, so an outage in one provider (say SMS) leaves the other three queues draining normally. User info, device info, and templates sit in cache because every one of 185 sends per second needs them and the database cannot take that read rate directly.

04

Reliability: never lost, sometimes twice


The hardest requirement is unstated: a notification may arrive late or twice, but it must never vanish. The fix is persistence plus retry. Every event is written to a notification log database first, and any send that fails goes back onto its queue for another attempt. Duplicates then become the price of safety. In a distributed design the send can succeed while its acknowledgement is lost, so the worker resends something the user already has. The chapter is blunt: exactly-once delivery is impossible here, and the reference material it cites explains why no protocol fixes that.

  1. On arrival, compare the event ID against delivered IDs and discard repeats.

  2. Persist the event in the notification log before any send is attempted.

  3. On send failure, return the event to its queue for a bounded number of retries.

  4. When retries are exhausted, alert developers instead of silently dropping the event.

Interview tipInterview line: when asked whether recipients get each notification exactly once, say no, then earn the point back. Most of the time delivery is exactly once, crashes cause the occasional duplicate, and the event ID check keeps duplicates rare. That answer shows you know both the mechanism and its limit, which is what the follow-up is testing.
05

The extras that make it production


Five additions turn the queued design into a production system, and each guards one numbered requirement. Notification templates keep the millions of daily sends in a consistent format with parameters and tracking links, which cuts errors and build time. The notification setting table of user, channel, and opt-in flag is checked before every send, which enforces requirement (4). Rate limiting caps how many notifications one user receives, because flooding inboxes makes users disable everything. App key and app secret pairs restrict the send APIs to verified clients, since open send APIs are a spam cannon. Queue-depth monitoring watches the backlog per queue and adds workers before delays breach requirement (2). Event tracking through an analytics service records open rate, click rate, and engagement across states such as pending, sent, delivered, clicked, and unsubscribed, which is how the team learns what to send less of.

The lines that matter below are the delivered ID set, the duplicate check before sending, and the requeue on failure. The set is the dedupe store from the previous section. The check runs before any provider call, so a redelivered event dies cheaply. The except branch returns the event to the queue, which is the retry loop from the flow diagram. A real system adds two things this sketch omits: a cap on attempts so a poison event cannot spin forever, and the developer alert when the cap is hit.

Show code · python
python
import time
from collections import deque

class FakeProvider:
    def __init__(self):
        self.attempts = {}
    def send(self, event):
        n = self.attempts.get(event['id'], 0) + 1
        self.attempts[event['id']] = n
        if event['id'] == 'sms-2' and n == 1:
            raise ConnectionError('provider timeout')
        return True

seen = set()  # delivered event IDs: the dedupe store
queue = deque([
    {'id': 'push-1', 'channel': 'push', 'to': 'device-token-9f3'},
    {'id': 'sms-2', 'channel': 'sms', 'to': '+65-8100-1234'},
    {'id': 'push-1', 'channel': 'push', 'to': 'device-token-9f3'},
])
provider = FakeProvider()
while queue:
    event = queue.popleft()
    if event['id'] in seen:
        print('drop duplicate:', event['id'])
        continue
    try:
        provider.send(event)
        seen.add(event['id'])
        print('delivered:', event['id'])
    except ConnectionError:
        print('failed, requeue:', event['id'])
        queue.append(event)
Q&A

Check yourself


Q1Your SMS provider goes down for an hour. Push and email providers are healthy. What happens to push and email?
  • Push and email keep flowing; only SMS waits in its own queue
  • All channels pause until the SMS provider recovers
  • Push and email are rerouted through the SMS queue
✓ Push and email keep flowing; only SMS waits in its own queue — Each channel owns its queue, so the SMS outage only blocks the SMS queue while push and email workers keep draining theirs.
Q2A worker crashes after sending but before recording success, and the user gets the same push twice. What removes the duplicate?
  • Persist every event twice so one copy always survives
  • Make the third-party service promise exactly-once delivery
  • Check the event ID against delivered IDs and drop repeats
✓ Check the event ID against delivered IDs and drop repeats — Crashes between send and acknowledgement cause resends, and exactly-once is impossible in a distributed setup, so the event ID dedupe is the real fix.
Q3Users start disabling all notifications after a flood of promotional pushes. What protects opt-in rates?
  • Add more workers so promotional bursts deliver faster
  • Cap sends per user per day and check opt-in before every send
  • Merge promotions into one daily digest for every user
✓ Cap sends per user per day and check opt-in before every send — Overmessaging drives users to switch everything off, so a per-user frequency cap plus the opt-in check protects requirement (4) directly.
Sources: System Design Interview Vol 1: Ch. 10, Design a notification system (pp. 151-165)