IM Systems in Depth

What happens betweenAna’s phone and Ben’s?

Design a messaging system step by step, from one server to a million people online, one decision per chapter. Server and client both, in English and Chinese.

Start with the whole picture

A messaging system does two things: receive every message safely, then deliver it to every device, including one that comes back after being offline.

Not drawn: voice and video calls, end-to-end encryption, multi-region failover, retention and compliance, search.

App backendbusiness systemsOS pushAPNs, FCMthird partyAnaStill on Saturday?the app’s SDK: outbox, local copy, resendsBenStill on Saturday?✓a phone and the web app, each syncsinternetIM serverBusiness (beside)called, not passedaccounts, authfriends, blockschecked before ACKgroups, membersopen APImessages, callbacksConnectionOne open connection per online device, so we can pushcost: online devices; stateful, a restart reconnects everyoneMessageReceives: no loss, no repeats, in order; replies “sent” once storedcost: messagesDispatchDelivers to every device: push if online, notify if notcost: messages × recipients (group of 500: 500 inbox writes, 750 pushes)StorageKeeps messages, inboxes and each device’s sync pointcost: messages × retention (× recipients with inbox copies)↑ receive: once stored, Ana gets “sent” and stops waiting↓ deliver: to every deviceaccountsif offlinewakemessage log1234
  1. Send: Ana’s message enters the IM over her connection
  2. Once stored, reply “sent” (not “delivered”): Ana waits for nothing after this
  3. Push to Ben’s online devices; offline ones get a notification, only a hint
  4. Any device that missed something pulls to catch up when it returns

Point at or tap a layer to see its job, why it is hard, and its chapters.

Written so far: the opening (chapter 0), first draft.

Every chapter goes in this order

  1. The problemA concrete scene, such as “the same message appeared twice”.
  2. Watch it failRun the simple design in a simulator on the page.
  3. EstimateA number you can check on paper explains why it fails.
  4. Fix itSwitch the fix on in the same simulator and compare.
  5. The costWhat the fix costs, and what other products choose.

Why this exists

There are many articles about messaging architecture. Most show a final diagram, or one company’s history. Few explain, step by step, why a system ends up the way it does.

Here, one chat app grows from 1,000 users to 1,000,000 people online. The four layers and the business layer beside them start crowded into one program, and each step of growth forces one of them out and deeper.

The last chapters change the product (toB workplace chat, community servers, live rooms, a private messenger, customer service, an IM platform for other businesses) and show the same decisions flipping. There is no single right architecture: the product decides.

Chapter map

The trunk follows one app from v0 to v3; each stop gives its scale and why it has to grow. Then a focus on reliability, and finally the branches change the product. The dot before each chapter is the layer it is mostly about, in the same colours as the diagrams.

connectionmessagedispatchstoragebusinessclientall layersBlue titles: published (first drafts for now)

  1. Opening

    See the whole system first.

    1. 0The whole picture
  2. v0: one server

    1,000 users

    “We want chat in our app.” Make it work first.

    1. 1The smallest chat
    2. 2Push instead of poll
  3. v1: correct on phones

    100,000 users

    On mobile networks, messages get lost, repeated and out of order.

    1. 3Acks and retries
    2. 4Duplicates
    3. 5Ordering
    4. 6Offline sync
    5. 7The local database
    6. 8Images and files
    7. 9v1: a reliable channel
  4. v2: groups and devices

    100,000 users, groups of 500

    Product adds groups of 500 and a desktop app.

    1. 10Fan-out
    2. 11Joining and leaving a group
    3. 12Who may message whom
    4. 13Unread counts and read receipts
    5. 14Many devices
    6. 15Waking an offline phone
    7. 16Recall and edit
    8. 17The conversation list
    9. 18v2: groups at 100,000 users
  5. v3: many servers

    1,000,000 online

    One server cannot hold them, and every deploy disconnects everyone.

    1. 19Split the gateway
    2. 20Presence
    3. 21A million connections
    4. 22Heartbeats
    5. 23Reconnect storms
    6. 24Accounts: the app’s users and the IM’s
    7. 25App lifecycle
    8. 26The web client
    9. 27Sharing out the conversations
    10. 28A queue before storage
    11. 29One SDK, many platforms
    12. 30Old apps, new servers
    13. 31v3: a million online
  6. Focus: reliability

    Once the system is big, the hard part is staying up: knowing messages arrive, and losing nothing through overload, deploys, failing dependencies and a lost data centre.

    1. 32Where did my message go?
    2. 33Overload: limits and degradation
    3. 34Deploys without disconnects
    4. 35When a dependency fails
    5. 36Losing a data centre
    6. 37Drills and postmortems
  7. Branches: other products

    The same system for toB, communities, live rooms, customer service, a platform: which decisions flip?

    1. 38Workplace chat (toB)
    2. 39Community servers
    3. 40Live rooms
    4. 41Private messenger
    5. 42Customer service
    6. 43IM as a platform
    7. 44Same decisions, different products

Side trips: on the wireoptional

Optional reading on the protocol.

  1. W1Framing
  2. W2Encoding
  3. W3TCP or UDP
  4. W4Group encryption