# Chapter 8: Images and files

> A photo is a thousand times a text message. Should it travel the same road?

IM Systems in Depth · https://im.liko.page/en/media-messages/

By chapter 7, text messages on the phone are no longer lost, repeated or out of order, and a killed app loses none of
them. But chats are not only text.

Ana climbs to the top of the mountain, sends Ben a photo, then three lines: “Made it!”, “So windy”, “Dinner when I’m
down?”. Up there she has one bar of signal.

Every mechanism so far was built for 200-byte messages. This chapter asks: **can a photo take the same road? If not,
which one?**

## 1. How big is a photo

Take a photo from the phone’s camera as **4 MB**; before sending, apps usually shrink and recompress it, say to
**200 KB**. Both numbers are this chapter’s assumptions; phones and apps differ a lot, but the order of magnitude holds.

- The compressed photo is 200,000 ÷ 200 = **1,000 times** a 200-byte text message; the original is **20,000 times**.
- Upload time on a weak network like the mountain top, assumed 2 Mbit/s up, that is 250 KB a second: the original takes
  **16 s**, the compressed one **0.8 s**, a text under 1 ms. At a better 8 Mbit/s: 4 s and 0.2 s.
- Side trip W1 capped one frame on the message connection at 64 KiB. The original is **61 times** that, the compressed
  photo **3.1 times**. Sending the original on the message connection means raising the cap to about 4 MiB; in W1’s worst
  case, with each of v1’s 10,000 connections stuck holding half a message, that is 10,000 × 4 MiB ≈ **41.9 GB** of
  buffers, against 0.66 GB at 64 KiB.

To be fair: with compressed photos and a good network, a photo holds the connection for 0.2 s, putting it in the message
connection works, and chapters 3 and 4 give it the outbox, ACKs, resends and dedup for free. The trouble is that people
also send originals, videos and files of tens of MB, and the signal is often poor. So look at the hard case: a 4 MB
original, a weak network, one drop in the middle.

## 2. Watch it fail

Ana sends the photo (the 4 MB original), then a line at 3, 7 and 10 s. The uplink carries 250 KB a second. At 12 s the
network drops for 2 s; with chapter 2’s 0.6 s handshake on top, the connection is back.

*[Interactive figure: open the page to use it, https://im.liko.page/en/media-messages/]*

The default is “Photo on the message connection”: the photo is one 4 MB message sharing the connection with the texts.

- **The texts queue behind the photo.** The bytes on one connection go in order (side trip W1: TCP is one byte stream,
  and one frame must finish before the next), and the outbox sends in order too. The three lines were sent at 3, 7 and
  10 s, but they wait for the photo. This is **head-of-line blocking**: one big item at the front holds up everyone
  behind it.
- **One drop, and the photo starts over.** By the drop, 250 KB/s × 12 s = **3 MB** has gone out, but half a message is
  useless to the server; when the connection is back, the outbox sends the 4 MB again from the first byte. The texts
  wait all over again.
- The result: the photo is stored at **30.7 s**, and the three lines wait **27.7, 23.7 and 20.7 s**; the phone sends
  3 + 4 = **7 MB** for this photo, 3 MB of it for nothing.

Switch to “Flaky” (a drop every 8 s): each time the connection comes back, 5.4 s remain, not enough for a 16 s photo.
The photo never finishes, not one line gets out, and in 60 s Ben receives nothing while the phone has sent **13.6 MB**.
On a train or in a lift, the user does not feel “the photo is slow” but “the whole chat is stuck”.

## 3. Estimate

**A drop at 70%: how much is sent again.** The original takes 16 s on the weak network; at 70%, 11.2 s in, with
2.8 MB sent, the link drops.

- **One request**: half a file is useless to the server; when the connection is back, the whole **4 MB, 16 s** goes
  again, and the 2.8 MB already sent is wasted.
- **In parts**: parts of 512 KiB (524,288 bytes, Telegram’s largest part), each kept by the server once complete. At the
  drop, 5 parts (2.62 MB) are complete and the 6th loses 179 KB. Back online, the phone asks “how far did you get?” (one
  round trip, 0.2 s) and sends the remaining **1.38 MB, 5.51 s**. Wherever the drop falls, it costs at most one part,
  **2.1 s** of uploading.

**When drops are frequent.** Say drops come at random, on average every T seconds; each one costs R = 2.6 s with the
reconnect, and then the upload starts over. A stretch of u seconds of sending gets through without a drop with
probability e^(−u/T), so on average it takes e^(u/T) tries; counting the time each failed try wastes plus R, it takes
(T + R)(e^(u/T) − 1) seconds on average (the standard result for a task that restarts on failure).

- T = 30 s: the 16 s original in one request takes **23.0 s** on average (the chance of no drop in 16 s is
  e^(−16/30) ≈ 59%); in parts of about 2 s each, **18.0 s**.
- T = 8 s: one request takes **67.7 s** on average, because the chance of no drop is only e^(−2) ≈ 14%, so nearly every
  try is wasted; in parts, **24.2 s**.

The shorter each piece, the likelier it gets through before the next drop.

**The bytes the server gains.** Assume 5% of v1’s 4,000,000 messages a day are photos (an assumption):

- 4,000,000 × 5% = **200,000 photos** a day, × 200 KB = **40 GB**. All the text in a day is 4,000,000 × 200 bytes =
  **0.8 GB**. Photos are 5% of the messages and **50 times** the bytes of the text.
- A year is **14.6 TB**, **43.8 TB** with the book’s 3 copies; text is 292 GB a year, 0.88 TB with 3 copies.
- At peak, 139 messages a second, 6.95 of them photos: **11.1 Mbit/s** of uploads. Downloads are larger: a message
  reaches 10 devices on average, and if each downloads the full photo, that is **111 Mbit/s**. Chapter 2 found that
  pushing all the text at peak takes **2.2 Mbit/s**.

A message is small, must be fast and in order; a photo is large, written once and never changed, and can be cached:
they belong in different places, rather than growing the message table 40 GB a day and keeping the message server’s
network card busy moving photos.

## 4. The fix: the file takes its own road

Sending a photo becomes a few steps:

1. **Compress.** The phone shrinks and recompresses the photo (4 MB → 200 KB), computes its SHA-256 hash, width and
   height, and makes a tiny placeholder. If the user ticked “original”, it skips the compression.
2. **Ask for an upload address.** On the message connection the phone asks: “I want to upload a 4 MB file.” The server
   runs the business checks (may Ana send files, is it under the size limit), assigns a file id and returns an upload
   address. The address points to the upload service below and carries a time-limited token signed by the server:
   whoever holds it can upload this one file there, without any other credentials. It is the same idea as S3’s
   [presigned URLs](https://docs.aws.amazon.com/AmazonS3/latest/userguide/PresignedUrlUploadObject.html).
3. **Upload in parts, resume after a drop.** The phone opens another connection and sends the file part by part. What
   remembers “which parts have arrived” is an **upload service** in front of object storage: once it has all the parts, it puts them together and stores the whole file. After a
   drop, the phone asks the upload service how far it got and carries on from there. This is a resumable upload; [tus](https://tus.io/protocols/resumable-upload) is a
   ready-made open protocol for it: `POST` creates an upload, `PATCH` with `Upload-Offset` writes into it, and after a
   drop a `HEAD` tells how much the server has.
4. **Send a small message.** With the file uploaded, the phone sends an ordinary message on the message connection,
   whose content is a reference to the file:

```json
{ "type": "image", "file_id": "f_7Kq2…", "size": 4000000, "sha256": "9f2c…",
  "width": 4032, "height": 3024, "mime": "image/jpeg", "blurhash": "LKO2?U%2Tw=w]~RBVZRi};RPxuwH" }
```

   About **300 bytes**, the size of a line of text. It takes the road chapters 3 to 7 fixed: outbox, ACK, resend, dedup,
   seq, sync, all of it. The message server checks that the file id exists and was uploaded by Ana. The size and hash
   can be trusted only if the server computed them itself: the upload service computes them once it has every part; or
   the presigned URL signs the size and SHA-256, and S3 [computes the checksum
   itself](https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity-upload.html) and rejects
   the upload if it does not match. Otherwise a client can lie about the hash.
5. **Ben downloads on demand.** The message reaches Ben’s phone. The bubble reserves space from the width and height and
   decodes the `blurhash` into a blurred placeholder ([BlurHash](https://blurha.sh/) describes an image’s rough colours
   in 20 to 30 characters); then it downloads a thumbnail of about 10 KB; only when Ben taps does the full photo come.
   Downloads go through a CDN: a finished file never changes, so the CDN and the phone can cache it freely.

In the simulator above, switch to the middle “Separate upload, one request”: each line takes only **0.1 s**, no longer
queued behind the photo; but the photo still starts over after the drop and is stored at **30.9 s**. Switch to
“Separate upload, in parts”: the photo is stored at **20.6 s**, 10.3 s sooner, with **4.18 MB** sent and only
**0.18 MB** wasted. On “Flaky”, the one-request photo never finishes in 60 s; the one in parts is through at **28.5 s**.

With no drops, the separate upload is a little slower than the message connection: the photo is stored at 17.1 s
instead of 16.1 s. The extra is the round trip to ask for the address (“address” in the figure), a handshake for the
upload connection, and the final small message: the fixed price of two roads, about 1 s a photo.

**What downloading on demand saves.** Assume 30% of recipients tap to see the full photo (another assumption). The
server’s downloads a day fall from 2,000,000 × 200 KB = **400 GB** with every device downloading the full photo to
2,000,000 × (10 KB + 30% × 200 KB) = **140 GB**, and at peak from 111 to **39 Mbit/s**, most of it carried by the CDN.
Ben receives 267 × 5% ≈ 13 photos a day; his downloads fall from **2,670 KB** to **935 KB**. All his text is 53 KB.

## 5. The outbox gains a step

Before, a message in the outbox had one step: send, wait for the ACK. A photo message has two: upload, then send. Ana’s
screen gains states:

- **Uploading 37%**: the photo sits where she sent it, with a progress bar;
- **Sending**: the file is up, the small message waits for its ACK;
- **✓ Sent**, or **! Failed**. A failure must say which step: if the upload failed, a retry continues the file; if the
  send failed, the file is already on the server and a retry sends only the 300-byte message.

These states live in chapter 7’s local database: the `messages` row’s `state` gains `'uploading'`, and an `uploads`
table keeps the local file’s path, its hash, the file id, the upload address and its expiry, and the number of parts
confirmed so far (`parts_done`).

Chapter 7’s rule still holds: each state change is committed in the same transaction as the fact it stands for. When
the file is up, setting `state` to `'sending'` and recording `file_id` are one transaction. `parts_done` is only the
phone’s own note; what counts is the upload service’s answer. If the app is killed halfway, on the next launch it finds
the `'uploading'` row, asks the upload service how far it got, and carries on. If the upload address has expired, it
asks for a new one.

**On the receiving side**, photos live in the app’s cache folder and the database keeps only the file id and the path;
photos come to about **487 MB** a year (**1.12 GB** if every one were downloaded in full) against 33.6 MB of text,
which is why chat apps have “clear cache”.

**Upload the same file only once?** A forward keeps the same file id; the server checks that the forwarder can see the
file, and not a byte is uploaded. Deduplicating by hash across **all users** (“a file with this hash is already here,
skip the upload”) also saves uploads, but it tells anyone whether someone else has uploaded a given file: a side
channel that leaks information ([Harnik et al., 2010](https://doi.org/10.1109/MSP.2010.187)). Whether and how to
deduplicate safely is a separate decision, not this chapter’s.

## 6. The cost

- **The two roads must agree.** With the file and the message apart, one can succeed while the other does not:
  - The file is up but the message never goes (the app was deleted, the user cancelled): object storage holds a file
    nobody refers to. It must be cleaned up, say, deleted if no message refers to it within a few days.
  - The message arrives but the file does not: Ben sees an image that will not open. So the phone must wait for the
    upload to be confirmed before sending the message, and the server must check the file exists when the message
    comes.
- **The order changes.** The photo message goes out only after the upload, so its seq (chapter 5) comes after the three
  lines sent meanwhile. Ana’s screen shows [photo, Made it!, So windy, …]; Ben sees [Made it!, So windy, …, photo]. Either
  move Ana’s photo down when the ACK comes, to match Ben, or accept that the two see different orders. Another way is to
  send a **placeholder message** first (“Ana is sending a photo”), which takes its seq the moment she taps send, and turn
  it into the real photo when the upload is done; that needs “editing a sent message”, chapter 16.
- **Who can download.** A file id or download address is itself the permission: whoever holds it can download. The
  usual answer is a signed, time-limited download address: the CDN checks the signature and expiry at the edge, then
  serves the same cached file, as with [CloudFront signed
  URLs](https://docs.aws.amazon.com/AmazonCloudFront/latest/DeveloperGuide/private-content-signed-urls.html). The cost is
  a coarse check: until it expires, anyone holding the address can download, and after Ben is removed from the
  conversation, the addresses he already has keep working until they expire. Matrix moved to [authenticated
  downloads](https://spec.matrix.org/latest/client-server-api/#content-repository) in version 1.11
  (`/_matrix/client/v1/media/download/…`) and deprecated the unauthenticated `/_matrix/media/v3/download`; but it
  checks only that the user is logged in, not that they are in the room.
- **More systems to run.** The upload service, object storage, the CDN, signing upload addresses, cleaning up expired
  files. Storage grows by 43.8 TB a year, so someone must decide how long photos are kept.

## 7. Other answers

- **Keep one connection, but cut the photo into small frames interleaved with the texts.** HTTP/2 does this: one
  connection carries many streams whose frames [can be interleaved](https://www.rfc-editor.org/rfc/rfc9113#section-5),
  so a large file does not block small requests. That removes head-of-line blocking and allows resuming. The cost: every
  photo byte still passes through the message server, the 40 GB a day and 111 Mbit/s peak included.
- **Telegram** uploads files part by part with MTProto’s `upload.saveFilePart`: per its
  [documentation](https://core.telegram.org/api/files), parts are at most 512 KiB, and such requests are best kept on
  separate sessions and connections that run nothing else. The same idea in its own protocol. A message carries an
  “extremely low-res thumbnail” (`photoStrippedSize`), which plays BlurHash’s role.
- **Matrix**: its [content repository](https://spec.matrix.org/latest/client-server-api/#content-repository) is a
  separate upload service. The client calls `POST /_matrix/media/v3/upload` and gets back an address of the form
  `mxc://<server-name>/<media-id>`; an `m.image` message carries only that address, plus width, height, type, size and a
  thumbnail in `info`.
- **Straight to S3**: S3’s multipart upload needs [parts of at least
  5 MiB](https://docs.aws.amazon.com/AmazonS3/latest/userguide/qfacts.html), and its documentation suggests multipart
  from about 100 MB; so a 4 MB photo sent straight to S3 is one `PUT` that starts over after a drop. Small resumable
  parts need your own upload service in front.

## 8. This chapter’s decision

*[Interactive figure: open the page to use it, https://im.liko.page/en/media-messages/]*

**Decision card**

- Problem: A photo is 1,000 times a text message (an original 20,000 times). Sent inside the message connection on a weak network, every text behind it waits; one drop starts the whole file over; on a flaky link it never finishes. Photo bytes are 50 times the text and would fill the message server’s disk and network.
- Choice: Files take their own road: the phone asks for an upload address and uploads the file in parts to an upload service in front of object storage, asking how far it got after a drop; only then does it send a message of about 300 bytes with the file id, size, hash, width × height and a placeholder. The receiver shows the placeholder and a thumbnail first and downloads the full photo from a CDN on tap. The outbox gains an “uploading” state; each step is committed with its state in the local database in one transaction.
- Cost: The two roads must agree: orphan files to clean up, and no message before its file; the photo message’s seq comes after texts sent during the upload; a file address is itself the permission, and a signed address cannot be taken back before it expires; an upload service, object storage, a CDN and 43.8 TB more storage a year to run.
- Revisit when: In a 500-member group one photo is seen by 750 devices, so downloading on demand matters more (v2’s capacity tables, chapter 18); whether photos follow a user to a new device (chapter 14); a placeholder message edited later (chapter 16); with end-to-end encryption the server cannot see file contents (chapter 41).
- Other answers: Small frames interleaved on one connection (as in HTTP/2; the bytes still go through the message server); Telegram’s upload.saveFilePart; Matrix’s content repository (mxc://); tus; S3 presigned URLs and multipart upload.

Every piece of v1 is now in place. The next chapter puts them together: it follows one message along the whole path,
switches the mechanisms of chapters 3–8 off one at a time and shows, with each chapter's own simulator, which failure
comes back; then it redoes the server's and the phone's accounts from v0 to v1.
