Chapter 3
Acks and retries
The phone says “sent”, yet the message never arrived. What went wrong?
v0 works, there are more users, and the complaints follow: “I saw it go out, and they say they never got it.”
In a lift, Ana sends Ben “I’m downstairs”. Her phone shows it sent; Ben never gets it. No error, no warning; the message is simply gone.
This chapter starts v1: correct on phones. v1 deals one by one with the old troubles of mobile networks: messages get lost (this chapter), repeated (chapter 4), out of order (chapter 5), devices that were offline must catch up (chapter 6), and the app must keep its own copy (chapter 7). Chapter 2 covered the receiving side (push is a hint, the pull is the truth); this chapter covers the sending side. First: what does “sent” actually mean?
1. How it works now
In chapter 1, Ana sent with a POST /messages request, and the server replied with the message’s id once it was in
the table: that reply was in fact a confirmation. With the long-lived connection of chapter 2, sending moved onto
that connection too, with no separate request. The code that is easiest to write then: write the message into the
connection, and if write() returns without an error, mark it “sent”. The old request’s reply is quietly gone.
Chapter 2 described this trap on the server’s side when pushing: a successful write() only means the bytes are in
the local send buffer. The same trap applies on the phone: the bytes are in the phone’s own buffer, which says
nothing about whether they left the phone, let alone whether the server stored them. If the connection drops then
(in a lift, in a tunnel, switching networks), whatever was in the buffer is gone, and the app knows nothing. Even
staying with POST, which does have a confirmation, there is no outbox and no automatic resend: when the request
times out or the app is killed, nobody knows whether it arrived, and nothing is left to send again.
A dropped connection is only one case. Between Ana pressing Send and the server storing the message, any step can fail:
- On the phone: the app is killed by the OS before sending, or the battery dies.
- On the network: the connection breaks midway: a lift, a tunnel, switching between Wi-Fi and 4G.
- On the server: an internal error; a database write that fails or times out; an overloaded server shedding requests it cannot handle; a restart for a deploy, landing exactly between “received” and “stored”.
These failures are all different, but to the phone they look the same: no “stored” comes back. So the fix below covers them all with one acknowledgment, instead of handling each kind on its own.
How the simulator models “lost”: TCP retransmits lost packets itself, so data on a live connection does not silently lose a piece; what really happens on phones is that the connection breaks with data still in flight. The simulator abstracts all of this as: each leg (the message going up, the ACK coming down) is lost with some probability. The series assumes 10%, far higher than real systems, on purpose: so the problem shows up within 8 messages. The slider goes all the way down to 0.
2. Watch it fail
With 10% loss, Ana sends 8. How many will her phone show as “sent” while the server never stored them?
a try×lost on this leg ✓shown sent ✓shown sent, never stored stored never stored
- Stored by the server
- –
- Shown sent, never stored
- –
- Extra copies stored
- –
- Shown failed
- –
- Tries in all
- –
In the default set, message 4 is lost on the way, yet Ana’s phone put a ✓ on it the moment it was written. The server never saw it, Ben will never get it, and Ana thinks he did.
This is worse than “failed to send”: a failed message gets resent by the user; a message shown as sent is never looked at again.
3. Estimate
With 10% lost on each leg:
- No acknowledgment: 10% of messages are lost on the way and still show as sent. In the long run, about one in ten “sent” marks is false. Over 10,000 simulated messages it is 10.1% (random variation).
- Small rates, real numbers: real systems lose far less than 10%, but v1 carries 4,000,000 messages a day. Losing just 0.1% means 4,000 a day shown as sent and never delivered; 0.01% is still 400. To each user, every one of them is “but I sent it”.
- Wait for the ACK: a try succeeds only if both legs do, so the server’s ACK gets back with 0.9 × 0.9 = 0.81. On average a message takes 1 ÷ 0.81 ≈ 1.23 tries.
- At most 5 tries: a try fails with 1 − 0.81 = 0.19, all five with 0.19⁵ ≈ 0.025%, about 2.5 in 10,000. Those honestly show “failed” and let the user decide whether to resend, instead of falsely showing “sent”.
- When to resend: not once a second, but waiting longer each time: in the simulator, tries leave at 0, 1, 3, 7 and 15 seconds, and after another 16 seconds without an ACK the app gives up, at 31 seconds.
4. The fix: an outbox, an ACK, resends
Switch the simulator above to “Wait for the ACK, resend without it”:
- Outbox: every message first goes into an outbox on the phone, and leaves it only when acknowledged. The outbox lives on disk, not in memory: after the OS kills the app or the phone restarts, unsent messages must still be there (chapter 7 covers the local database; chapter 1’s “in practice” said that memory is lost on restart).
- Acknowledgment (ACKACKACK(确认)The receiver’s “got it” to the sender. In this book, the server’s ACK means the message is in the message log, not that anyone has received it.See the glossary): the server sends it only after it has stored the message. As chapter 0 said, a message counts as received once it is written and stored, and that is when the ACK goes out; so “sent” means “the server stored it”, not “the other person has it”. What exactly “stored” means is refined in chapter 28. If the server knows something went wrong (a failed write, overload), it should say so with an error rather than stay silent: for a temporary error the phone retries later; for a permanent one (rejected, too large, blocked) it marks the message failed at once and stops. The timeout is only the last resort, for when no answer arrives at all.
- Resend: no ACK after a while, send again by the schedule above. If the connection is still up, resend on it; but a missing ACK usually means the connection is gone, and then “resend” means reconnecting first and sending every unacknowledged message in the outbox again; the backoff schedule really paces the reconnects. The simulator leaves reconnecting out and uses a timer per message instead. Resending does not keep the order either: if 5 is acknowledged while 4 is still being resent, 4 ends up after 5 (chapter 5). The 1-second timeout is there so the simulator is easy to follow; on a real mobile network one round trip can exceed a second when the signal is weak, and timeouts are usually several seconds, or the app resends needlessly and makes more duplicates. Doubling the wait only slows retries down; millions of phones that lost the network at the same moment will still come back in the same second if they all follow the same schedule, so each wait also gets some random jitter to spread them out (AWS’s post explains it well; chapter 23 is about reconnect storms).
- Failed: in the simulator, five failed tries mark a message failed. Real apps are usually more patient: with no network they keep it “sending”, count retries only while connected, and give up after a time (a minute or so) rather than only a number of tries, so a 40-second lift ride does not turn messages red. And remember that “failed” only means “no ACK came back”: the message may have been stored after all (say the last ACK was lost), and if the user taps resend there is one more copy, which chapter 4 handles.
- On screen: sending → ✓ sent → ! failed (tap to resend).
The same 8 messages are now all stored, and no “sent” is false. Message 4 was lost on its first try and arrived when resent a second later.
5. The cost: duplicates
But the server row now has a red ×2. Message 7 arrived on its first try, but the server’s ACK was lost on the way;
the phone thought it had not arrived, sent it again a second later, and the server stored it twice.
- On a first try, “the message arrives, its ACK is lost” happens with 0.9 × 0.1 = 9%; a resend can again arrive and lose its ACK, so in all about 10% of messages are stored more than once, about 11 extra copies per 100 messages (some are stored three times). Over 10,000 simulated messages: exactly 10% and 11.1.
- Ben sees the same line two or three times. Chapter 4 fixes it with a message ID: the server recognises “I have seen this one”, ACKs it again, and does not store it again.
One small cost besides: a message that needs a resend waits at least one timeout longer.
6. Other answers
- “Doesn’t TCP already retransmit? Why acknowledge ourselves?” TCP only covers a live connection. When the connection breaks or the app is killed, whatever was in TCP’s buffers is gone; and even a TCP acknowledgment only means the server’s kernel got the bytes, not that the server’s program stored the message. Only the application knows what it meant to send and whether it was stored. Reliability has to be ensured at the two ends: the end-to-end argument.
- Telegram’s MTProto: acknowledgments are messages themselves
(
msgs_ack); client and server acknowledge to each other which messages they received. - XMPP stream management (XEP-0198): both sides count what they send and acknowledge it, so after a reconnect they know what to send again.
7. This chapter’s decision
Next chapter: Ben sees the same line twice.