# Chapter 2: From polling to push

> The app reaches 100,000 users and messages should arrive at once. Why not just ask more often?

IM Systems in Depth · https://im.liko.page/en/push-or-poll/

The app has grown to 100,000 daily users, 10,000 online at peak. People start saying chat “feels slow”: a message
takes over a second to show up on the other phone. The product manager wants messages to arrive at once.

The obvious move is to ask more often: every 1 second, every 0.5 seconds, instead of every 2.

## 1. What asking more often does

Chapter 1 worked out that polling waits half the interval on average, plus two one-way network legs (0.2 seconds
in all): wait − 0.2 s = T ÷ 2. And 10,000 phones each asking every T seconds make 10,000 ÷ T requests a second.
Multiply the two and T cancels out:

> requests per second × (average wait − 0.2 s) = phones online ÷ 2 = 5,000

Wait and load are tied together: halve the wait and you double the requests. Drag the blue point in the chart below:

*[Interactive figure: open the page to use it, https://im.liko.page/en/push-or-poll/]*

- Every 2 seconds: **5,000 requests a second**, 1.2 seconds’ wait on average.
- Every second: 10,000 a second, 0.7 seconds.
- Every 0.2 seconds: **50,000 a second**, 0.3 seconds.

Yet at peak only about 1,389 deliveries a second are actually needed (139 messages a second, each reaching 10
devices on average in the series’ assumptions). So at least 72%, 86% and 97% of those requests come back empty.

To be fair: one decent server and database can take five or ten thousand small indexed queries a second. The real
problems are two:

- **Wait can only be bought with load**, and each step costs more. To wait a little less, you pay double again.
- **The phones pay too**: with the app open, a request every second or two spends data and battery on empty
  answers.

To get out of this trade you have to turn it around: instead of the phone asking, the server tells the phone when
there is something.

## 2. How the answer evolved

From “the phone asks” to “the server tells” there were a few steps. Each solved one problem and left another:

1. **Short polling** (chapter 1): works everywhere; most requests are empty, and a message waits half the interval
   on average.
2. **Long polling**: the phone sends a request and the server does not answer yet; it parks the request until a
   message arrives or about 30 seconds pass, then answers, and the phone sends the next request at once. A message
   can be answered as soon as it arrives, so the wait is close to the network time; requests are at most about
   1,389 + 10,000 ÷ 30 ≈ **1,722 a second** (one answer may carry several messages, so fewer in practice),
   far below polling every 2 seconds. The 30-second timeout is deliberate: common load balancers and proxies cut
   connections that stay silent for about 60 seconds. Between an answer and the next request there is a short gap,
   and a message arriving then waits for the next request; because the phone asks with `after=<id>`, it is only
   late, never lost.
3. **HTTP streaming (SSE, Server-Sent Events)**: an answer that never ends; the server writes each message into
   it. It only goes from server to client, so the phone still sends with separate requests. It has reconnecting
   built in: the server tags each event with an `id`, and after a drop the browser reconnects with the last one it
   got (`Last-Event-ID`), which is exactly `after=<id>`. The most common use today is AI chat streaming its answers
   ([Anthropic](https://platform.claude.com/docs/en/build-with-claude/streaming) and
   [OpenAI](https://platform.openai.com/docs/api-reference/streaming) both use SSE): one question, one long answer,
   then done. Messaging needs both directions, and a connection that stays.
4. **WebSocket** (native apps often use a custom protocol over TCP instead): it starts as an ordinary HTTP request
   carrying `Upgrade: websocket`; the server answers `101` to agree, and from then on the TCP connection stops
   speaking HTTP and carries frames (a frame is roughly one message): a two-way pipe that stays open. The server can
   write to the phone at any time, without being asked, and each message costs a few bytes of frame header instead
   of a request’s few hundred bytes of HTTP headers. Most messaging systems today use this kind: Discord’s clients
   connect to its [Gateway](https://docs.discord.com/developers/events/gateway) over WebSocket, and Telegram’s apps
   keep a long-lived connection speaking [MTProto](https://core.telegram.org/mtproto).

*[Interactive figure: open the page to use it, https://im.liko.page/en/push-or-poll/]*

**WebSocket or native TCP?** Two ways of doing the same thing, and often both at once:

- **WebSocket**: the only choice in a browser (besides SSE and long polling above); it rides on HTTP’s
  infrastructure, so port 443, load balancers, CDNs and proxies mostly let it through, and libraries are
  everywhere. The cost is one HTTP upgrade when connecting and a few bytes of frame header per message.
- **Native TCP with a custom protocol** (also wrapped in TLS): native apps only. It skips the HTTP upgrade, can
  use a tighter frame format, and leaves heartbeats, compression and encryption up to you. The cost is that you cut
  the byte stream into messages yourself (the side trip “Framing”), load balancers can only forward TCP, and some
  networks block unusual ports.

So the usual answer is one message protocol over two transports: apps on native TCP, the web on WebSocket, both
into the same connection layer. Telegram’s MTProto lists both
[TCP and WebSocket as transports](https://core.telegram.org/mtproto/transports).

## 3. Push

Take the last one: when a phone connects, it keeps one long-lived
connection (a WebSocket) to the server.

- The server keeps a table in memory: **user → their connections**. A set, not one: Ana has a phone and a laptop,
  and the series assumes 1.5 devices per person. This is the earliest form of chapter 0’s “online routes”.
- Since the connection stays open, Ana sends over it too, with no separate request.
- Once Ana’s message is stored, the server finds each of Ben’s connections and writes the message into it.
- The wait becomes the network time, about 0.2 seconds; about 1,389 frames a second, none of them empty.

In the chart above, push is the green square at the bottom left, far below the polling curve: it is outside the
“load for wait” trade. The two timelines under the chart are the same 6 messages: polling waits 1.6 seconds on
average, push 0.2.

**But keep the pull.** Under those two timelines, tick “Ben’s phone drops off from 6 s to 14 s”: the pushes sent
while it was off are lost, because the connection is gone. When the phone reconnects, it first pulls with
`after=<id>`, and everything it missed comes back.

Why can a push be lost? A successful write on the server only means the bytes are in the server’s own send buffer,
not that the phone has them. If the connection has died quietly (common in a lift or when switching networks), the
server takes a while to notice, and whatever it wrote meanwhile is gone. So: **push is a hint; the pull is the
truth.** The phone drops repeats by id. And the cursor (chapter 1’s `after`) only moves forward on what a pull
returns: pushes from different senders do not always arrive in id order. Say 8 is pushed first and 7 is still on
its way; if the phone moved its cursor to 8 on seeing it, and 7 then got lost in a drop, the reconnect pull would ask
for “after 8” and 7 would never come back. Pulling is not airtight yet either: as chapter 1 said, a smaller id may
commit later. Chapters 5 and 6 make all of this exact with a gap-free sequence per conversation.

## 4. What setting up a connection costs

That “one long-lived connection” is almost always encrypted: `wss://` (WebSocket over TLS) on the web, a custom
protocol over TLS in apps. Unencrypted `ws://` is not just unsafe; proxies along the way often break it. So a
connection takes a few steps first:

| Step | Round trips | On mobile (0.2 s per round trip) |
|---|---|---|
| TCP three-way handshake | 1 | 0.2 s |
| TLS handshake (certificate, key exchange) | TLS 1.3: 1; TLS 1.2: 2 | 0.2 s / 0.4 s |
| HTTP upgrade to WebSocket | 1 | 0.2 s |
| **Total** | **3–4** | **0.6–0.8 s** |

(DNS lookup not included.) So a phone needs most of a second just to connect before it can receive its first push,
and after that each message takes 0.1 seconds one way. That is the point of a long-lived connection: **you pay this
once**, not for every message.

Connecting has a hidden bill too: **server CPU**. A TLS handshake does public-key work (signing with the
certificate, exchanging keys), far more expensive than encrypting ordinary data afterwards. The series assumes one
server does about 2,000 full TLS handshakes a second. That is deliberately conservative, roughly a few CPU cores
signing with RSA-2048; a many-core machine with an ECDSA certificate does ten times more. The conclusions below
hold either way:

- **A restart brings everyone back at once**: 10,000 phones ÷ 2,000 a second = **5 seconds** of the server doing
  nothing but handshakes, delaying new messages too. At v3, with 100,000 connections on one gateway, it is 50
  seconds (chapter 23, reconnect storms).
- **Polling does not escape this bill**: if every poll opened a new HTTPS connection, polling every 2 seconds would
  mean 5,000 handshakes a second, 2.5 times what one server can do. So polling also relies on HTTP keep-alive to
  reuse connections, and keep-alive is already a “somewhat long” connection.

Ways to save: TLS session resumption (session tickets, TLS 1.3’s PSK), which lets a reconnect skip the most
expensive step, certificate signing and checking; and not letting everyone come back in the same second (chapter
23).

There is also a newer path: **HTTP/3**, over QUIC on UDP, which merges the transport and TLS 1.3 handshakes into
one, so a new connection takes **1 round trip** (TCP + TLS 1.3 take 2); a resumed connection can even send data
with no round trip at all (0-RTT, which TLS 1.3 has too), though such data can be replayed, so it is only safe for
requests that are harmless to repeat, and sending a message is not one. WebSocket over HTTP/3 is still rare, and
browsers often fall back to TCP; what to do when UDP is blocked and how to use connection migration are left to the
side trip “TCP or UDP” and chapter 25.

## 5. The cost

- **The server now keeps state.** Long polling already did (parked requests and a table of who is waiting; in
  event-driven servers such as Go or Node a parked request holds no thread, while a thread-per-request server
  cannot park many), and push makes it plain: the server remembers where everyone is connected. In the series’
  assumptions an idle connection takes about 30 KB of memory (stack, read and write buffers, TLS state, kernel),
  so 10,000 take **about 300 MB**. One server holds that; by chapter 21, with a million connections, it does not.
- **A restart drops everyone**, and they all reconnect together (see the handshake bill above; chapters 19 and 23).
- **Everything on the path must allow long-lived connections.** Load balancers, proxies and firewalls must let a
  connection stay open; mobile networks and phone OSes close connections that are silent for too long (heartbeats
  in chapter 22; what happens when the app goes to the background in chapter 25).

## 6. Where polling still lives

Long-lived connections do not win everywhere.

**In practice**: the author’s first chat system used WebSocket, but one business scenario, on a TV box or a similar
device as the author remembers it, could not open a WebSocket and fell back to long polling. Some client
environments simply cannot keep a WebSocket, and long polling is the fallback that works everywhere.

There are public examples too:

- **Bots**: Telegram’s Bot API offers two ways to get messages,
  [`getUpdates`](https://core.telegram.org/bots/api#getupdates) (long polling) and webhooks. A bot without a public
  HTTPS address simply long-polls.
- **The first step of real-time libraries**: [Socket.IO](https://socket.io/docs/v4/how-it-works/) connects with HTTP
  long polling by default and then upgrades to WebSocket, because some proxies and firewalls block WebSocket.
- **Phones in the background**: once the app is in the background its connection is closed; an OS push wakes it,
  and it pulls when opened, chapter 1’s “anything new?” again (chapters 15 and 25).

Polling survives for a few reasons: it gets through everything; the server keeps no state, scales easily and can
cache; the data is not urgent; or the client simply cannot keep a connection.

## 7. The v0 summary

*[Interactive figure: open the page to use it, https://im.liko.page/en/push-or-poll/]*

**All of v0**: one server. The connection layer is now its own part, holding long-lived connections and the
user → connections table; the message and dispatch layers are still in the same program, and the business layer is
still a few checks; one database.

**How much it holds**: 10,000 connections and 139 messages a second at peak: one server is enough. Here is the bill
for the whole machine and for each phone (polling counted with HTTP/1.1 and keep-alive):

**The server (10,000 online)**

| Resource | Polling every 2 s | Push |
|---|---|---|
| CPU | 5,000 requests a second, parsing HTTP and querying | about 1,389 frames a second; cheap normally, handshakes in a burst at restart |
| Memory | keep-alive also holds about 10,000 idle connections, but they can close any time, and no user → connection table | about 30 KB per connection, about 300 MB for 10,000, plus the user → connection table |
| Network bandwidth | a few hundred bytes to 1 KB of HTTP headers on each request and reply: 5,000 × 2 × 0.5–1 KB ≈ 40–80 Mbit/s, almost all headers | the messages themselves: 1,389 × 200 bytes ≈ 2.2 Mbit/s |
| NIC (packets) | a request, a reply and the ACKs, about 3–4 packets each: about 15,000–20,000 packets a second | a frame and its ACK, about 2 packets: about 2,800 packets a second |
| Disk space | messages about 0.8 GB a day; with one 200-byte access-log line per request, about 3.6 GB per peak hour, about 29 GB a day at average load (about a third of peak) | messages about 0.8 GB a day; far less logging |
| Disk IO | writes: about 139 inserts a second at peak, usually one fsync (forced flush to disk) per commit; reads: 5,000 queries a second, random reads on every cache miss; logs about 1 MB a second appended | writes: the same 139 a second; reads: only a pull per reconnect |

A NIC counts packets, not bytes: with many small packets, what fills up first is often the NIC and the CPU handling
receive interrupts. Today’s NICs usually handle millions of packets a second, far beyond v0; the bill matters once
connections reach the millions across many gateways (chapter 21). Also, each connection counts as an open file (a
file descriptor). Linux’s default soft limit is still often 1,024 per process, and going past it fails with “too
many open files”; Go raises it itself since 1.19, otherwise it is a line in ulimit or systemd, and memory, above, is
the real limit. HTTP/2 and HTTP/3 compress repeated headers, so the bandwidth row shrinks a lot, but the request
count, the empty answers and the phone wake-ups below do not change.

**In practice**: in the author’s experience, logs written for business needs often made disk IO the bottleneck,
before the reading and writing of messages themselves. The 1 MB a second in the table is only one access-log line
per request, which any disk takes easily; business logs are often several lines per request, each longer, written
synchronously in small pieces, on the same disk as the database. Log volume follows the request count, and polling
makes several times the requests of push. The usual fixes: log only on the key paths and sample the many repeated
requests; buffer logs in memory and write them in batches, to a separate disk, or ship them to other machines.

**One phone (one hour with the app open)**

| Resource | Polling every 2 s | Push |
|---|---|---|
| Requests | 1,800 | one push per message received |
| Data | HTTP/1.1 headers alone: 1,800 × 2 × 0.5–1 KB ≈ 1.8–3.6 MB | the messages themselves, about 200 bytes each |
| Battery | wakes the cellular radio every 2 seconds, and it stays in high power about 10 seconds after each wake-up, so it never sleeps | wakes it once per incoming message |
| Background | the OS will not let a background app ask every 2 seconds | the OS closes the connection; an OS push wakes the app (chapters 15 and 25) |

How much battery push saves depends on how busy the chats are. In the series’ assumptions each online device
receives about 1,389 ÷ 10,000 ≈ 0.14 messages a second at peak, about 500 an hour, one every 7 seconds; that busy,
push keeps the radio awake about three quarters of the time, not far from polling’s always. In a quiet hour (say 20
messages), push wakes it 20 times, about 3 minutes, while polling still keeps it awake the whole hour. So push saves
most when chats are quiet; when they are busy, the answer is to batch several messages into one send (chapters 22
and 25). Long-lived connections will also need heartbeats (chapter 22).

Push wins on network, disk, and the phone’s battery and data when chats are quiet; it loses on the server having to
remember where everyone is connected, and on the CPU burst at restart.

**What breaks next**: **connections break.** In a lift or on a train, a connection can drop in the middle of a send,
and the phone cannot tell whether Ana’s message reached the server. Sending it again may make two; two messages may
also arrive in different orders. That is v1 (chapters 3–9); chapter 3 models these cases as packet loss.

## 8. This chapter’s decision

**Decision card**

- Problem: 10,000 phones online and users want messages at once; polling can only trade load for wait.
- Choice: One WebSocket long-lived connection per device; the server pushes once a message is stored; on connect and reconnect the phone pulls by after=<id>, and push is only a hint.
- Cost: The server keeps state (user → connections, about 30 KB each); 3–4 round trips to connect and TLS handshakes cost CPU, and a restart brings everyone back at once; everything on the path must allow long-lived connections.
- Revisit when: When connections outgrow one server and deploy reconnects become too much (chapters 19, 21 and 23); heartbeats and the background (chapters 22 and 25); when networks get worse and messages are lost or repeated (chapter 3 on).
- Other answers: Long polling (the fallback where WebSocket is blocked); SSE (when only the server needs to push).
