Ten million pushes, five million emails, and one million texts leave one system every day, and none may be lost. This lesson builds that system: scope the load, meet each delivery channel, break the single server into queues and workers, then add the retry and guardrails that make it production.
Outcomes
The interview opens with scope questions, and the answers fix four numbered requirements. (1) Three channels: mobile push, SMS, and email. (2) Soft real time: a user should get each notification as soon as possible, but a slight delay under heavy load is acceptable. (3) Volume: 10 million pushes, 1 million SMS, and 5 million emails per day. (4) Respect opt-out: a user who opts out receives nothing further. Triggers come from client applications or from server-side schedules, and supported devices span iOS, Android, and laptop or desktop.
Run the envelope live. 16 million a day divided by 86,400 seconds is about 185 sends per second on average: roughly 116 push, 58 email, and 12 SMS. That rate is only the floor. Each send also pays for rendering content, a database or cache lookup for contact info, and a round trip wait on a third-party service. One server absorbing 185 of those compound sends per second, with peaks above it, is the load that breaks the first design in the next section.
Each channel delivers through different hands. An iOS push needs three parts: a provider that builds the request with a device token and a payload, Apple Push Notification Service which relays it, and the device itself. The payload is a JSON dictionary with an alert title and body plus a badge count, as in Bob asking you to play chess. Android follows the same path through Firebase Cloud Messaging instead of APNs. SMS goes through commercial gateways such as Twilio or Nexmo. Email usually goes through services such as Sendgrid or Mailchimp rather than a self-built mail server, for better delivery rates and analytics.
Contact info lives in two simplified tables. The user table holds email addresses and phone numbers; the device table holds device tokens, and one user can own several devices, so one push can fan out to every device that user logged in on. Extensibility matters here because providers differ by market: FCM is unavailable in China, where alternatives such as Jpush and PushY take its place. A design that hardcodes one provider per channel cannot enter those markets.
The first design is one notification server between triggering services and the third-party services, and it fails in three numbered ways. (1) Single point of failure: one server down means no channel sends anything. (2) Hard to scale: database, cache, and per-channel processing cannot grow independently inside one box. (3) Performance bottleneck: HTML rendering plus waiting on third-party responses burns the one server fastest exactly at peak hours. Requirement (3) is what turns these from theory into an outage.
How to read: Step through the six stages in order, then follow the retry edge from Workers back to the queues and ask what happens when the third-party service is down.
A triggering service calls the send API on the notification servers
Servers validate addresses, fetch contact info from cache or DB, and enqueue
The event waits in its own type queue: iOS, Android, SMS, or email
A worker pulls the event and calls the matching third-party service
On failure the event returns to the queue for retry instead of being lost
The improved design moves the database and cache out of the server, runs many notification servers behind autoscaling, and puts a message queue between stages. Servers now do only four jobs: expose internal APIs to triggering services, validate emails and phone numbers, fetch rendering data from cache or DB, and enqueue events. Workers pull events and call third parties. Each notification type owns its queue, so an outage in one provider (say SMS) leaves the other three queues draining normally. User info, device info, and templates sit in cache because every one of 185 sends per second needs them and the database cannot take that read rate directly.
The hardest requirement is unstated: a notification may arrive late or twice, but it must never vanish. The fix is persistence plus retry. Every event is written to a notification log database first, and any send that fails goes back onto its queue for another attempt. Duplicates then become the price of safety. In a distributed design the send can succeed while its acknowledgement is lost, so the worker resends something the user already has. The chapter is blunt: exactly-once delivery is impossible here, and the reference material it cites explains why no protocol fixes that.
On arrival, compare the event ID against delivered IDs and discard repeats.
Persist the event in the notification log before any send is attempted.
On send failure, return the event to its queue for a bounded number of retries.
When retries are exhausted, alert developers instead of silently dropping the event.
Five additions turn the queued design into a production system, and each guards one numbered requirement. Notification templates keep the millions of daily sends in a consistent format with parameters and tracking links, which cuts errors and build time. The notification setting table of user, channel, and opt-in flag is checked before every send, which enforces requirement (4). Rate limiting caps how many notifications one user receives, because flooding inboxes makes users disable everything. App key and app secret pairs restrict the send APIs to verified clients, since open send APIs are a spam cannon. Queue-depth monitoring watches the backlog per queue and adds workers before delays breach requirement (2). Event tracking through an analytics service records open rate, click rate, and engagement across states such as pending, sent, delivered, clicked, and unsubscribed, which is how the team learns what to send less of.
The lines that matter below are the delivered ID set, the duplicate check before sending, and the requeue on failure. The set is the dedupe store from the previous section. The check runs before any provider call, so a redelivered event dies cheaply. The except branch returns the event to the queue, which is the retry loop from the flow diagram. A real system adds two things this sketch omits: a cap on attempts so a poison event cannot spin forever, and the developer alert when the cap is hit.
import time
from collections import deque
class FakeProvider:
def __init__(self):
self.attempts = {}
def send(self, event):
n = self.attempts.get(event['id'], 0) + 1
self.attempts[event['id']] = n
if event['id'] == 'sms-2' and n == 1:
raise ConnectionError('provider timeout')
return True
seen = set() # delivered event IDs: the dedupe store
queue = deque([
{'id': 'push-1', 'channel': 'push', 'to': 'device-token-9f3'},
{'id': 'sms-2', 'channel': 'sms', 'to': '+65-8100-1234'},
{'id': 'push-1', 'channel': 'push', 'to': 'device-token-9f3'},
])
provider = FakeProvider()
while queue:
event = queue.popleft()
if event['id'] in seen:
print('drop duplicate:', event['id'])
continue
try:
provider.send(event)
seen.add(event['id'])
print('delivered:', event['id'])
except ConnectionError:
print('failed, requeue:', event['id'])
queue.append(event)