IM Systems in Depth

Side trip W2

Encoding

How much bigger is a message in JSON than in binary? Write 300 as a varint (AC 02).

This chapter is a first draft. It will be revised once the first eight chapters are written.

Side trip W1 put a length in front of every message and cut the byte stream into frames. How the bytes inside a frame are laid out is still open.

As of chapter 4, a message the server pushes to Ben carries these things: the id the server gave it when storing it (chapter 1), the message IDmessage ID消息 IDA unique ID the client gives each message before sending it, unchanged on resends. The server uses it to recognise a resend.See the glossary msg_id the phone made before sending (chapter 4, a 128-bit random number), the conversation conversation_id, the sender sender_id, the text text, and the time the server got it, created_at (milliseconds). The most common way to write it is JSON:

{"id":482913207,"msg_id":"744ea99e-2638-aecd-4b0a-119e631d2d54","conversation_id":1204387,"sender_id":300,"text":"I’m downstairs","created_at":1791198000412}

All spaces are already gone, and msg_id is the usual UUID string. The other common way is Protobuf: first give every field a number in a schemaschemaschema(消息定义)One file that says which fields each message has, their types and their numbers, from which the server’s and every app’s code is generated. Protobuf’s .proto files are one kind.See the glossary, then put only numbers and values on the wire, never names.

message Message {
  uint64 id = 1;
  bytes  msg_id = 2;           // 16 bytes
  uint64 conversation_id = 3;
  uint64 sender_id = 4;
  string text = 5;
  uint64 created_at = 6;       // milliseconds
}

To be clear from the start: nothing breaks in this side trip. JSON is a perfectly respectable choice: Slack’s Socket Mode hands apps their events as JSON, Discord’s gateway can speak JSON, and Matrix is JSON from end to end. The question is where the two differ, and whether the difference is worth a switch.

1. One message, two sets of bytes

Guess first: Ana’s “I’m downstairs”, as the server pushes it to Ben. Written as JSON, how many times the size of the same message in Protobuf?

Pick a field to see its bytes on both sides

abfield name / tag and length12value’one cell per byte: the apostrophe is 3 bytes, so 3 cells wide

JSON
159 bytes
Protobuf
–
JSON ÷ Protobuf
–

JSON is 159 bytes, Protobuf 56: 2.8 times. Pick the fields one by one and see where the bytes go:

  • Field names. JSON writes "conversation_id": and the rest into every message: the 6 names with their quotes and colons are 64 bytes, plus 7 for braces and commas. Protobuf writes one byte per field, a tag (the field’s number and type; section 2 has the details).
  • Numbers. JSON writes decimal characters, a byte per digit: created_at is 13 bytes. Protobuf uses a varintvarintvarint(变长整数)Small numbers in few bytes: 7 bits a byte, the top bit meaning “more follows”. 0–127 take one byte; 300 is AC 02. Protobuf writes integers and field numbers this way.See the glossary (small numbers take fewer bytes; section 2 again): 6 bytes for the same number.
  • msg_id. 16 random bytes. As a UUID in JSON they are 36 characters plus quotes; Protobuf puts the 16 raw bytes, plus 2 bytes of tag and length. (Unpadded base64 in JSON would be 22 characters, 14 bytes fewer.)
  • The text. The same on both sides: the same 16 UTF-8 bytes (the curly apostrophe in “I’m” takes 3).

So the longer the text, the smaller the ratio. With a 60-byte text it is 203 to 100 bytes, 2.0×; with 150 bytes, 293 to 191, 1.5×. (That is about 60 and 150 English letters, or 20 and 50 Chinese characters at 3 bytes each in UTF-8.)

2. Writing 300 as a varint

Ana was the app’s 300th user, so her sender_id is 300. In Protobuf this field is three bytes: 20 ac 02.

The tag is field number << 3 | type. sender_id is field 4 and type 0 is a varint: 4 << 3 | 0 = 32 = 0x20. Type 2 means “a length follows”, for strings and bytes, so the tag of text (field 5) is 5 << 3 | 2 = 0x2a. Field numbers 1 to 15 fit their tag in one byte, and Protobuf’s guide suggests keeping them for the most frequent fields.

The value: a varint holds 7 bits per byte, low bits first; a top bit of 1 means more bytes follow.

  1. 300 in binary is 1 0010 1100: 9 bits, too many for the 7 bits of one byte.
  2. Take the low 7 bits, 010 1100 = 0x2c; more follow, so set the top bit: 0x2c | 0x80 = ac.
  3. What is left, 300 >> 7 = 2, fits, with a top bit of 0: 02.

So 300 is ac 02. 0 to 127 take one byte, up to 16,383 two; the server id 482,913,207 takes 5 bytes and the millisecond timestamp 6. (Negative numbers have their own encoding, zigzag; this message has none.) The encoding guide tells the same story with 150 → 96 01.

3. Estimate: what those bytes are worth

Compress first. WebSocket has a standard extension, permessage-deflate; browsers ask for it, and it is used if the server turns it on. By default it shares the compression context across a connection: a message can refer back to bytes in earlier messages, so the field names JSON repeats over and over cost almost nothing.

Take 50 messages in one conversation (a batch assumed here: Ana and Ben taking turns, all different short texts, in Chinese). On average, per message:

Bytes per message JSON Protobuf Ratio
Raw 160.8 58.3 2.76
Each alone 144.6 62.2 2.32
Shared context 70.7 47.7 1.48
Whole page 52.7 43.2 1.22
  • Raw: the server has compression off. This row is a page average: inside a page each Protobuf message costs 2 more bytes for its tag and length (JSON only a comma), and half the messages are Ben’s, whose id is longer, hence 2.76×, not 2.8×.
  • Each alone: every message compressed from scratch (no_context_takeover was agreed). Short messages barely shrink, and the binary ones grow (58.3 → 62.2), which is why servers usually compress only messages above some size.
  • Shared context: messages pushed one by one, with the compressor kept. This is the row for everyday chat: messages pushed one at a time to Ben while he is online.
  • Whole page: all 50 in one message, compressed together. This is the row for a phone that comes back and fetches a batch of missed messages at once.

Compression eats JSON’s bulk; most of what remains is the UUID’s hex characters and the decimal digits.

The server: at v1’s peak, 1,390 deliveries a second, each 159 − 56 = 103 bytes larger, about 143 KB a second: 1.15 Mbit/s. Nothing, for one server; and for a phone, a hundred-odd bytes more per message is nothing either.

Compression has a cost too: a shared context means keeping compression state for every connection. By zlib’s own figures, with default settings the compressor for the server’s outgoing messages takes about 268 KB; if the phone compresses its messages too, the server also keeps a decompressor, about 40 KB. That is about 308 KB per connection, while the series counts 30 KB for an idle connection; v1’s 10,000 connections would need 3.08 GB. RFC 7692 lets the ends agree on a smaller window (256 bytes to 32 KB), but that shrinks only the window: with a 1 KB window a connection still needs about 150 KB, because the compressor’s 131 KB of hash table and output buffer is set by the server’s own memLevel, not by the window (lowering it also compresses worse).

How much CPU parsing takes depends on the language and the library; there is no public figure to borrow here, so this side trip does not estimate it.

The conclusion: at v1’s scale, bytes are not a reason to change formats. Uncompressed, the gap is 2.8×, but that bandwidth is worth little; compressed, it is 1.2–1.5×, at the price of a few hundred KB of memory per connection.

4. The real reason: old and new versions live together for years

Chapter 5 will add a sequence number (seq)seq序号(seq)The number of each message in a conversation, assigned in order by the message service that owns it. It sets the order, lets a client see which message is missing, and tells a returning phone which conversations it is behind on.See the glossary to messages. From that day the new server puts seq into every message, but the apps on people’s phones do not update that day, and some do not update for a year. Turn on “Add chapter 5’s seq” in the scene: JSON grows by "seq":1284, 11 bytes; Protobuf by 38 84 0a, 3 bytes. What does an old app do with that message?

  • Protobuf: 0x38 is 7 << 3 | 0. The old code has no field 7, but type 0 says a varint follows, so it reads it and moves on. Skipping unknown fields is part of the format, the official runtimes for every language do it, and proto3 even keeps them and writes them back out, so they survive being passed along.
  • JSON: it depends on each platform’s parser. Most ignore unknown keys by default: Swift’s Codable, Gson and Moshi on Android, Go’s encoding/json, and the browser’s JSON.parse takes anything. Some are strict by default: kotlinx.serialization defaults to ignoreUnknownKeys = false, and Java’s Jackson has FAIL_ON_UNKNOWN_PROPERTIES on by default; both throw on a key they do not know. One strict platform in one version is enough: the day a field is added, it cannot read messages. JSON can skip new fields too; it just relies on every platform remembering to be configured that way. It is a convention, not a rule of the format.

A few more points, all about “many versions, many platforms”:

  • Names and numbers are separate. Only numbers are on the wire, so a field can be renamed freely. The one rule: never reuse a number. Put deleted fields under reserved so no one later gives the same number another meaning; the guide lists what reuse leads to: parse errors, corrupted data, even leaked data.
  • One schema generates every platform’s code: Go, Swift, Kotlin and TypeScript get their types and encoding code from the same .proto, checked when they compile, instead of four hand-written copies.
  • 64-bit integers: JavaScript numbers are exact only up to 2^53 (about 9×10^15). v1’s ids are far below that; only with 64-bit, Snowflake-style IDs does JSON have to write them as strings, as Discord does.

To be fair, JSON can also have a schema (JSON Schema, OpenAPI) with generated code. So the real decision is “one schema with numbered fields”; once you have it, binary on the wire comes almost free, and with it 2.8× fewer bytes without keeping a compressor for every connection.

5. The cost

  • You can no longer read it. A packet capture or a log line shows 08 b7 d7 a2 e6 01 12 10 …. Without the schema, protoc --decode_raw can only say “field 1 is a varint”. So you need tools: logs print messages through Protobuf’s JSON mapping, and debugging proxies need the schema. Note that the JSON mapping’s parser rejects unknown fields by default, the opposite of the binary format; turn on “ignore unknown fields” when you use it.
  • One more repository, one more build step. The schema needs one place every platform shares; every platform generates code from it when it builds, and the versions must match. The web client carries one more library.
  • Rules to keep. Numbers are never reused, types are not changed at will (turn a uint64 into a string and old readers cannot read it); someone has to review schema changes.
  • Unfriendly to outsiders. An open API or a webhook for third parties should be JSON they can read directly.

6. Other answers

  • Telegram has its own TL serialization: a schema and binary as well, with integers a fixed 4 or 8 bytes instead of varints. It spends a few bytes on alignment in return for simple arithmetic; and only Telegram’s own tools understand it.
  • Signal uses Protobuf: the server’s Envelope and libsignal’s SignalMessage are defined in .proto files, with the costs of section 5.
  • Discord’s gateway takes JSON or ETF (Erlang’s built-in binary format), with zlib-stream or zstd-stream compression, one context per connection. It pays compression memory on every connection in return for JSON anyone can read.
  • Matrix is JSON throughout, with canonical JSON (one way to write each object) where it signs. It pays in bytes so that anyone can write a server or client in any language.
  • MQTT only carries bytes; what the payload looks like is up to the application, so this decision is still yours. Its own header writes lengths with the same 7-bits-a-byte variable-length integer.
  • CBOR (RFC 8949) and MessagePack (msgpack.org) are “binary JSON”: shorter numbers, but usually written as maps with field names, and no schema built in (CBOR can add one with CDDL). What they lack is compatibility rules; you have to agree on them yourself.

7. The decision

The frames are cut, and what goes inside them is settled. The main path goes on at chapter 5, “Ordering”: Ana and Ben send a line at almost the same moment, and the two of them see the lines in different orders.