# Chapter 3: Acks and retries

> The phone says “sent”, yet the message never arrived. What went wrong?

IM Systems in Depth · https://im.liko.page/en/acks-and-retries/

v0 works, there are more users, and the complaints follow: “I saw it go out, and they say they never got it.”

In a lift, Ana sends Ben “I’m downstairs”. Her phone shows it sent; Ben never gets it. No error, no warning; the
message is simply gone.

This chapter starts v1: **correct on phones**. v1 deals one by one with the old troubles of mobile networks:
messages get lost (this chapter), repeated (chapter 4), out of order (chapter 5), devices that were offline must
catch up (chapter 6), and the app must keep its own copy (chapter 7). Chapter 2 covered the receiving side (push is a
hint, the pull is the truth); this chapter covers the sending side. First: what does “sent” actually mean?

## 1. How it works now

In chapter 1, Ana sent with a `POST /messages` request, and the server replied with the message’s `id` once it was in
the table: that reply was in fact a confirmation. With the long-lived connection of chapter 2, sending moved onto
that connection too, with no separate request. The code that is easiest to write then: write the message into the
connection, and if `write()` returns without an error, mark it “sent”. The old request’s reply is quietly gone.

Chapter 2 described this trap on the server’s side when pushing: a successful `write()` only means the bytes are in
the local send buffer. The same trap applies on the phone: the bytes are in **the phone’s own** buffer, which says
nothing about whether they left the phone, let alone whether the server stored them. If the connection drops then
(in a lift, in a tunnel, switching networks), whatever was in the buffer is gone, and the app knows nothing. Even
staying with `POST`, which does have a confirmation, there is no outbox and no automatic resend: when the request
times out or the app is killed, nobody knows whether it arrived, and nothing is left to send again.

A dropped connection is only one case. Between Ana pressing Send and the server storing the message, any step can
fail:

- **On the phone**: the app is killed by the OS before sending, or the battery dies.
- **On the network**: the connection breaks midway: a lift, a tunnel, switching between Wi-Fi and 4G.
- **On the server**: an internal error; a database write that fails or times out; an overloaded server shedding
  requests it cannot handle; a restart for a deploy, landing exactly between “received” and “stored”.

These failures are all different, but to the phone they look the same: no “stored” comes back. So the fix below
covers them all with one acknowledgment, instead of handling each kind on its own.

How the simulator models “lost”: TCP retransmits lost packets itself, so data on a live connection does not silently
lose a piece; what really happens on phones is that **the connection breaks with data still in flight**. The
simulator abstracts all of this as: each leg (the message going up, the ACK coming down) is lost with some
probability. The series assumes 10%, far higher than real systems, on purpose: so the problem shows up within 8
messages. The slider goes all the way down to 0.

## 2. Watch it fail

*[Interactive figure: open the page to use it, https://im.liko.page/en/acks-and-retries/]*

In the default set, message 4 is lost on the way, yet Ana’s phone put a ✓ on it the moment it was written. The server
never saw it, Ben will never get it, and Ana thinks he did.

This is worse than “failed to send”: a failed message gets resent by the user; a message shown as sent is never
looked at again.

## 3. Estimate

With 10% lost on each leg:

- **No acknowledgment**: 10% of messages are lost on the way and still show as sent. In the long run, **about one in
  ten “sent” marks is false**. Over 10,000 simulated messages it is 10.1% (random variation).
- **Small rates, real numbers**: real systems lose far less than 10%, but v1 carries 4,000,000 messages a day. Losing
  just 0.1% means **4,000 a day** shown as sent and never delivered; 0.01% is still 400. To each user, every one of
  them is “but I sent it”.
- **Wait for the ACK**: a try succeeds only if both legs do, so the server’s ACK gets back with 0.9 × 0.9 = **0.81**.
  On average a message takes 1 ÷ 0.81 ≈ **1.23 tries**.
- **At most 5 tries**: a try fails with 1 − 0.81 = 0.19, all five with 0.19⁵ ≈ **0.025%**, about 2.5 in 10,000.
  Those honestly show “failed” and let the user decide whether to resend, instead of falsely showing “sent”.
- **When to resend**: not once a second, but waiting longer each time: in the simulator, tries leave at 0, 1, 3, 7
  and 15 seconds, and after another 16 seconds without an ACK the app gives up, at 31 seconds.

## 4. The fix: an outbox, an ACK, resends

Switch the simulator above to “Wait for the ACK, resend without it”:

- **Outbox**: every message first goes into an outbox on the phone, and leaves it only when acknowledged. The outbox
  lives on disk, not in memory: after the OS kills the app or the phone restarts, unsent messages must still be there
  (chapter 7 covers the local database; chapter 1’s “in practice” said that memory is lost on restart).
- **Acknowledgment (ACK)**: the server sends it only **after it has stored the
  message**. As chapter 0 said, a message counts as received once it is written and stored, and that is when the ACK
  goes out; so “sent” means “the server stored it”, not “the other person has it”. What exactly “stored” means is
  refined in chapter 28. If the server knows something went wrong (a failed write, overload), it should say so with an
  error rather than stay silent: for a temporary error the phone retries later; for a permanent one (rejected, too
  large, blocked) it marks the message failed at once and stops. The timeout is only the last resort, for when no
  answer arrives at all.
- **Resend**: no ACK after a while, send again by the schedule above. If the connection is still up, resend on it;
  but a missing ACK usually means the connection is gone, and then “resend” means reconnecting first and sending every
  unacknowledged message in the outbox again; the backoff schedule really paces the reconnects. The simulator leaves
  reconnecting out and uses a timer per message instead. Resending does not keep the order either: if 5 is
  acknowledged while 4 is still being resent, 4 ends up after 5 (chapter 5). The 1-second timeout is there so the
  simulator is easy to follow; on a real mobile network one round trip can exceed a second when the signal is weak,
  and timeouts are usually several seconds, or the app resends needlessly and makes more duplicates. Doubling the
  wait only **slows** retries down; millions of phones that lost the network at the same moment will still come back
  in the same second if they all follow the same schedule, so each wait also gets some **random jitter** to spread
  them out (AWS’s [post](https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/) explains it
  well; chapter 23 is about reconnect storms).
- **Failed**: in the simulator, five failed tries mark a message failed. Real apps are usually more patient: with no
  network they keep it “sending”, count retries only while connected, and give up after a time (a minute or so) rather
  than only a number of tries, so a 40-second lift ride does not turn messages red. And remember that “failed” only
  means “no ACK came back”: the message may have been stored after all (say the last ACK was lost), and if the user
  taps resend there is one more copy, which chapter 4 handles.
- **On screen**: sending → ✓ sent → ! failed (tap to resend).

The same 8 messages are now all stored, and no “sent” is false. Message 4 was lost on its first try and arrived when
resent a second later.

## 5. The cost: duplicates

But the server row now has a red `×2`. Message 7 arrived on its first try, but the server’s ACK was lost on the way;
the phone thought it had not arrived, sent it again a second later, and the server stored it twice.

- On a first try, “the message arrives, its ACK is lost” happens with 0.9 × 0.1 = **9%**; a resend can again arrive
  and lose its ACK, so in all about **10%** of messages are stored more than once, about 11 extra copies per 100
  messages (some are stored three times). Over 10,000 simulated messages: exactly 10% and 11.1.
- Ben sees the same line two or three times. **Chapter 4** fixes it with a message ID: the server recognises “I have
  seen this one”, ACKs it again, and does not store it again.

One small cost besides: a message that needs a resend waits at least one timeout longer.

## 6. Other answers

- **“Doesn’t TCP already retransmit? Why acknowledge ourselves?”** TCP only covers a live connection. When the
  connection breaks or the app is killed, whatever was in TCP’s buffers is gone; and even a TCP acknowledgment only
  means the server’s **kernel** got the bytes, not that the server’s program stored the message. Only the application
  knows what it meant to send and whether it was stored. Reliability has to be ensured at the two ends: the
  end-to-end argument.
- **Telegram’s MTProto**: acknowledgments are messages themselves
  ([`msgs_ack`](https://core.telegram.org/mtproto/service_messages_about_messages)); client and server acknowledge to
  each other which messages they received.
- **XMPP stream management ([XEP-0198](https://xmpp.org/extensions/xep-0198.html))**: both sides count what they
  send and acknowledge it, so after a reconnect they know what to send again.

## 7. This chapter’s decision

*[Interactive figure: open the page to use it, https://im.liko.page/en/acks-and-retries/]*

**Decision card**

- Problem: On mobile networks, a message can be lost on the way while the phone shows it sent.
- Choice: Messages go into an outbox first; the server ACKs once a message is stored; without an ACK the phone reconnects and resends with backoff (and jitter); only when it keeps failing is the message marked failed.
- Cost: A lost ACK means a duplicate (about 10% of messages stored more than once); resends add delay; the outbox must live on disk; “failed” only means no ACK came back.
- Revisit when: Duplicates (chapter 4); what “stored” means (chapter 28); reconnect storms when the network comes back (chapter 23).
- Other answers: Rely on TCP alone (not enough); MTProto’s msgs_ack; XMPP’s XEP-0198.

Next chapter: Ben sees the same line twice.
