Adam Absorption: How Modal Ships 500 MB Instead of 500 GB for RL Rollouts
Disclaimer: This blog written by AI 🤖
A frontier-scale checkpoint is around 500 GB. Shipping one to a rollout fleet in another region takes minutes to hours and kills any hope of weight updates landing in seconds. Nan Jiang’s AI Engineer talk argues you can send roughly 500 MB instead and have the rollout engine reconstruct a bitwise-identical weights version. The bet is that fewer than 1% of rollout-visible weights actually change between consecutive versions—and the reason is not gradient sparsity. Gradients are dense: about 99% of parameters get a nonzero gradient, and the FP32 master update is dense too. It is just small.
The mechanism is a small Adam step meeting finite precision. The rollout engine serves a BF16 view whose rounding boundary sits near theta over 256—about 0.0039 for a weight around 1—while a typical Adam step at RL post-training learning rates runs around three millionths, more than a thousand times too small to cross it. The master weights move and the served value does not; Jiang calls this Adam absorption. Lower-precision serving makes it sharper still: an internal run serving GLM 4.7 Air in FP8 saw 0.15% of weights change on the first step and settle near 0.05%. Once a lossless patch is the unit of synchronization instead of a checkpoint, the rollout fleet stops needing to live in the trainer’s cluster. Training keeps its all-reduce and fast fabric; rollout islands scatter across whatever regions and providers have GPUs right now.
Modal’s implementation is called Stitch: a sidecar makes any engine version aware of incoming patches, and scattered inference capacity becomes one elastic rollout fleet. The architecture Jiang describes is a bulletin board where the trainer posts patches and rollout workers pull what they need—500 GB down to 500 MB without sacrificing correctness. Open questions remain around Muon optimizers and fully async RL, but the core insight is infrastructural: RL post-training wants four things at once (fast training, fresh weights, geographic flexibility, and cost efficiency), and the wrong sync primitive—full checkpoints—forces you to sacrifice three of them.